Confidential inference on Routstr without trusted hardware

Illustration: a glowing sealed envelope with a padlock seal rides a conveyor through a dark relay station; the operator in the foreground cannot see inside; a distant data-center tower waits at the far end.

Written by

in

Short version: a Routstr node can now resell AI inference without being able to read your prompts or the answers, and without handing you its API key. It doesn’t need enclaves or any other special hardware. It is experimental, so read the caveats at the end. Paper, code, how to try it.

The problem: the middleman sees everything

Routstr is a marketplace for AI inference. You pay a node with Cashu ecash or Lightning, without an account or a credit card, and the node forwards your request to a provider such as Venice using its own API key. It is a reseller, and a useful one: you get pay-per-request access with good privacy towards the provider, because the provider only sees the node.

The catch is that the node sits in the middle and sees everything in plain text, including any file you paste into the conversation. It could log it, sell it or change it. It could also swap the model you asked for for a cheaper one and still bill you for the expensive one. Today you simply trust the node not to do that.

For many uses that is fine. For source code, contracts, health questions or anything you would not post publicly, it is a lot of trust to give to an anonymous operator you found on Nostr.

The usual answer: trusted hardware

The standard fix is a trusted execution environment (TEE), an enclave inside the processor that even the machine’s owner can’t look into. The node runs inside it, and you check a “remote attestation”, a signed statement from the chip vendor that says which software is running inside. I wrote about hardware attestation for private AI before.

It works if you do the checking. An attestation tells you that some code is running in an enclave, and it protects you only if you know what that code is. You have to verify the whole software stack inside, ideally rebuild it yourself and compare the measurements, and you have to do it again every time the operator ships an update. Almost nobody does this. People see a green checkmark and move on, and the trust shifts from the operator to the hardware vendor and to whoever built the image. Enclaves also have a long history of side-channel attacks.

Illustration: on the left, a tall, leaning tower of system layers (a chip, a server, gears, an operating system window, a shield, a terminal) with a person on a ladder inspecting it with a magnifying glass. On the right, a small cube on a flat base engraved with mathematical curves, next to a padlock and a certificate.
With an enclave you keep auditing a stack of software. We wanted something that only needs math and the certificates you already trust.

Our approach: let the node carry a sealed envelope

We started from a different question: could the node do its job without ever having the plaintext?

Your browser does something like this every day. When you open a website over HTTPS, your Wi-Fi router and your ISP carry the traffic but can’t read it, because your browser sets up an encrypted TLS connection directly with the website and checks its certificate. The router only relays it.

In confidential mode the Routstr node becomes that router. Your client (routstrd, the Routstr daemon on your machine) opens a TLS connection directly to the provider, through the node, and the node passes the encrypted bytes back and forth. It adds the line with its own API key, because that is how the provider knows who pays. Your prompt and the model’s answer are encrypted with keys the node does not have.

What you rely on is the provider’s ordinary TLS certificate, the same one your browser checks, plus math: zero-knowledge proofs and a two-party computation. You don’t need a special chip, and there is no attestation to re-check after every update.

What the node can’t do any more

  • It can’t read your prompt or the answer. It only ever handles encrypted data.
  • It also can’t change anything, in either direction. Your client is the TLS client of the provider, and the node never holds the keys for your request or for the answer. Every TLS record carries a MAC, a cryptographic checksum that only someone with those keys can compute, so the node can’t forge one. That means it can’t touch the model’s output at all: the text of the answer, the tool calls the model makes and their arguments, anything else in the response. If it flipped a single byte, your client would detect it and throw the answer away. The same goes for what you send: your messages, your system prompt and the tool results you pass back to the model. A modified record fails the provider’s check and the provider rejects it. The node can cut the connection, which is a denial of service, but it can’t alter what goes through. It just relays sealed records. (The provider itself still sees the plaintext, more on that below.)
  • It can’t swap the model or raise the token limit. The model name and the limits sit at the end of your encrypted request. The node holds back that last piece until your client proves, with a zero-knowledge proof, that it names exactly the model and limits you both agreed on. Until then the provider has an incomplete request and does nothing.
  • It can’t inflate the bill. The charge is computed from the usage numbers the provider itself reports at the end of the answer, at the prices the node signed before the request. Your client proves those numbers to the node without revealing the answer, and gets a signed receipt.

A zero-knowledge proof here is a way for one side to convince the other that a statement about encrypted data is true (“this request ends with the agreed model name”, “these are the usage numbers the provider sent”) without showing the data itself. The node learns one bit: yes or no.

And you can’t steal the node’s API key

The node has a secret too: its API key with the provider. If you are the one holding the TLS keys, could you decrypt the line where the node wrote its key?

In the first version of the protocol you could, if you could also see the node’s own network traffic, for example by tapping its uplink or colluding with its hosting provider. That is a reasonable assumption for some nodes, but a node operator shouldn’t have to bet on it.

The second version, which is what runs now, removes that condition. Your client and the node compute the TLS keys together, in a two-party computation, and neither side ever holds the full secret. You end up with the keys for your request and for the answer. The node ends up with the key for the one piece it writes, its API key line. Even if you recorded every byte of the node’s traffic, you would not get its key. If either side cheats during the joint computation, the other side notices before any key is used and the session stops.

Illustration: two hands from opposite sides each hold one half of a glowing key above a stream of data flowing to a server; neither holds the whole key.
Client and node compute the TLS keys together. Each ends up with only the part it needs.

What it costs

The joint key computation needs a lot of precomputed material for every request, and it has to be ready before your prompt goes out. The node does the heavy part of that preparation and your client mostly downloads it: about 22 MB down and about 1.2 MB up per request. An earlier version had the client upload about 21 MB, and on home connections that upload was the slow part.

I measured it against the live node routstr.cypherpunk.today from a good home connection (about 100 Mbit/s down). A cold request, with nothing prepared in advance, took about 14 to 16 seconds to the first token, and 5 to 12 seconds of that was receiving the precomputed material. On a slow or congested connection a cold request can take about a minute.

The heavy part does not depend on your prompt, and routstrd prepares it ahead of time, but only while you are using it. After a confidential request, it prepares the next session in the background, and the next request skips that step. While requests keep coming, it keeps one or two sessions ready (more during bursts). After a few idle minutes it closes them, which means an idle routstrd keeps nothing open and nothing reserved. A prepared session doesn’t connect to the upstream provider (Venice, for my node) until a request uses it. Pre-warming is on by default in my routstrd branch, which is what the setup guide below runs.

With a ready session, the slow part used to be the two zero-knowledge proofs, about 8 seconds. Their setup is now prepared ahead of time as well (one prepared proof channel per session), and the proofs on the request path take about 0.05 to 0.25 seconds. From the same home connection, a request that finds a prepared session gets its first token in about 1 second, and most of that is the provider’s own latency. Pre-warming also hides a slow connection, because the download happens before your next request instead of during it.

The browser version (more on it below) can show a timeline of each request. Here is a prepared one on my node:

Timeline of a prepared confidential request: setup, 2PC preprocessing (9.49 s) and proof channel setup (4.56 s) before the send line; after it the TLS handshake with 2PC phases A and B, the proofs, the provider's first token at 3.66 s, 15.5 s of streaming and settlement.
One prepared request in the browser demo, on routstr.cypherpunk.today. The blue line is when I pressed send, the red dashed line is the first token.

Time runs left to right. The blue line is the moment I pressed send, and everything to the left of it happened before that: the node checked my credential and reserved the cost (471 ms), and the two-party preparation ran for about 9.5 seconds, with the proof channel set up alongside it (4.6 s). All of that ran while I was still typing, which is what pre-warming means.

After I pressed send, the browser and the node computed the TLS keys together in two rounds (2PC phases A and B, 925 and 940 ms) as part of the handshake with the provider, which took 2.07 s in total. The node proved with π_NK (282 ms) that its credential record matches the published template, wrote that record with its API key (59 ms), and the provider started answering. π_C2S, the proof that the request ends with the agreed model and token cap, took 72 ms. The first token arrived at 3.66 s.

The proofs themselves take tens to a few hundred milliseconds. Most of the rest is the provider generating text: this answer streamed for 15.5 s. At the end, π_C3S (149 ms) discloses only the usage numbers, and the node signs a receipt. The whole settlement took 390 ms.

You pay for the inference like for any other request to that node. The node reserves the maximum possible cost up front and charges what the provider reported.

On top of that, a node may charge fees for confidential mode. They are part of its signed offer, so your client sees them and accepts them before anything is charged. A client that doesn’t accept them is refused and pays nothing. There can be a fee per completed request, and a fee for a session that ends without a completed request, for example a prepared session you never use, or one you abort. The second fee is what protects the node from people making it do the expensive preparation for free. If you use the session, you don’t pay it.

If the provider refuses the request before doing any work (a 4xx error or a 503), you pay only the per-request fee, if the node charges one. Other server errors (500, 502, 504) can cost up to the reservation, because the provider may already have billed the work. If your complete request never reaches the provider, you pay only the unused-session fee, if there is one. My node (routstr.cypherpunk.today) charges no per-request fee and 2 sats for a session that ends unused. A typical message costs a fraction of a sat.

What it doesn’t hide

The provider still sees your prompt, as it always has, so pick a provider whose privacy policy you trust (my node uses Venice, which says it does not store prompts). The node still sees which provider and model you use and when, how long your request is, and the sizes of the streamed chunks, which can leak a little about the answer. It can still refuse service or misbehave with your balance. That is ordinary business trust, and node reputation covers it as it does today. And nothing here proves which weights the provider actually runs; for that you rely on the provider.

Combining it with enclave providers

Earlier in this post I was critical of TEEs. My complaint is about where they sit: you shouldn’t have to trust, and keep re-verifying, the software of a reseller in the middle. For the provider, which sees your prompt anyway, attestation is a reasonable extra layer, and this design can be combined with it.

Our protocol covers the Routstr node. It can’t read or change your data, it can’t swap the model or the token cap, and it bills from the usage the provider reports. Because your client is the TLS endpoint and talks to the provider through the node, it can also check the provider’s own TEE attestation end to end, for example with confidential-inference providers such as PPQ, Tinfoil or Phala, or on Venice’s end-to-end encrypted (E2EE) model path. The node only relays sealed records, so it can’t interfere with that check either.

So you could have both: guarantees from math about the reseller in the middle, and the provider’s enclave attestation if you also want to limit what the provider itself can see. The design supports this, because the node is transparent to your client’s TLS session, but routstrd doesn’t do it yet.

Status: new and experimental

This is fresh work, in a testing phase. It is not merged into the upstream Routstr projects yet; the changes are open-source as branches on my github. The two-party computation is new cryptography that nobody outside has reviewed yet. Please don’t use it for anything where a bug would hurt you.

My node, routstr.cypherpunk.today, uses Venice as its provider, and that is what I tested end to end against the live API. The design doesn’t depend on Venice, and it should work with most OpenAI-compatible API providers. The provider needs standard TLS 1.3 with a normal WebPKI certificate and the cipher and curve the protocol uses (AES-128-GCM and P-256). It also has to handle requests predictably: wait for the whole request before it starts, let the last duplicate key in the JSON win (so the model and token cap pinned at the end of the body take effect), enforce the token cap, and report token usage in the response stream. A node operator checks this with a test tool before offering a provider, and the paper lists the exact requirements. For now it covers text chat completions, from routstrd or from a browser (see below).

The first page of the paper Confidential and verifiable upstream inference over standard TLS 1.3, by Juraj Bednar with coauthors Claude Opus 5.5, Claude Sonnet 5.5, Kimi K3, DeepSeek V4.1 Flash, GPT-6 Astra and GLM-5.3 Flash: the title, the authors, the abstract and the start of the table of contents.
The first page of the paper (PDF).

Try it

If you want to play with it, the setup guide walks you through running my routstrd branch locally, with the matching SDK branch and the prover binaries, against my node routstr.cypherpunk.today, which runs the confidential mode. The guide also covers running your own node.

In the browser

There is now also a browser version: a static web app that does the same confidential inference entirely in the page. The two-party computation and the zero-knowledge proofs run in WebAssembly, the page talks to the node over WebSockets, and it verifies the provider’s TLS certificate itself.

You can try it at jooray.github.io/routstr-confidential-pwa. Fund it with Lightning or a Cashu token, or paste a key you already have. It shows the node’s fees before you pay, and for every message what it cost and the provider’s certificate chain. On this node a prepared session you never use costs 2 sats, and a typical message costs a fraction of a sat. In my measurements, a cold request took about 12 to 14 seconds to the first token, and a prepared one about 3 seconds.

The confidential chat demo v0.1.10 in a browser: a markdown answer from Qwen 3.8 Flash marked verified, warm, first token 3.7 s, 0.461 sat. The side panel shows the provider host api.venice.ai, TLS 1.3, the certificate chain to GTS Root R4, and the proofs with their timings.
The demo (v0.1.10) on routstr.cypherpunk.today, with a verified answer from Qwen 3.8 Flash. The side panel shows what the browser checked itself: the provider’s certificate chain up to a bundled Mozilla root, TLS 1.3 with the key share computed jointly in the 2PC, and each proof with its timing. This answer cost 0.461 sat.

The library behind it, routstr-confidential-web, can be used in your own web app. The demo’s source is in routstr-confidential-pwa.

In a browser you trust the page you load, so the demo shows the exact commit it was built from, and the code is public. Keys and balances live in the browser, so keep them small.

Please send feedback and bug reports, and try to break it. I would much rather someone finds the holes now than after people start relying on it.