Written at the proof-of-concept stage. Several things it lists as not yet built have shipped since: BYOK, OIDC single sign-on, the hash-chained WORM audit trail and Vault HA. The threat model has the current status.
Sepulchre: a credential broker that can't read your credentials
An engineering deep-dive into building zero-knowledge by policy, not by promise.
The problem nobody admits they have
Every company moves secrets by hand. A vendor needs your production API keys, so someone pastes them into an email. A contractor needs a database password, so it goes over Slack. A new hire needs the signing certificate, so it lands in a shared doc that outlives three reorgs. We have vaults and secret managers for secrets at rest, and we have TLS for secrets in transit between machines — but the moment a secret has to pass between two people, we fall back to the most insecure channels we own, and we do it constantly.
The usual "solution" is a one-time-secret pasteboard: a web app where you paste a secret, get a link, and the recipient views it once. These are useful, and they all share the same fatal asterisk: the service operator can read everything. The secret sits in their database, or their Redis, or their memory, in plaintext or under a key they hold. You are trusting the pasteboard company exactly as much as you were trusting email — you've just added a third party to the list of people who can read your production keys.
Sepulchre is an attempt to remove that asterisk. It is a credential broker whose central, load-bearing property is that the operator running it cannot read the credentials that flow through it — and, crucially, can prove they can't, with a policy file you can diff and a script you can run against the live system.
This is the story of how it's built, why the design choices are what they are, and what it taught us — including the bug that only surfaced when we drove the whole thing end to end.
Zero-knowledge by policy
"Zero-knowledge" is an overloaded phrase. In cryptography it means a specific class of proofs. In marketing it usually means "we encrypt things and pinky-swear." Sepulchre means something narrower and more honest, which we call zero-knowledge by policy:
No operator-scoped identity holds any capability that permits reading the path where submitted credentials live.
The enforcement isn't a clever protocol or a homegrown crypto scheme. It's a HashiCorp Vault policy with one decisive line:
path "sepulchre/data/{{TENANT}}/inbound/*" { # rendered per tenant
capabilities = ["deny"]
}
Submitted credentials live only in Vault KV under sepulchre/data/<tenant>/inbound/<intake-id>. The
operator — the human running the console, or anyone who steals their session — can list the
metadata of those submissions (so the dashboard can show "received / pending") but carries an
explicit deny on the secret data. Vault evaluates deny before allow, so no amount of
policy stacking, token wrapping, or capability union can override it. The only identity class
permitted to read inbound data is a service account authenticating with an AppRole token —
the machine that actually needs the credential, never a person.
Why is this better than "we encrypt it"? Because it's falsifiable. The policy is published
in the repository as the canonical artifact. A script, tools/verify-policy.sh, reads the
live policy out of the running Vault, normalizes it, and diffs it against the published file;
then it mints a real operator-scoped token and tries to read an inbound path. If that read
succeeds, the script exits non-zero and the build is broken. The guarantee is proven by the
read being denied, not by anyone reading the policy and nodding. CI runs this on every push
against an ephemeral Vault, and — because a guard you've only ever seen pass is a guard you
don't trust — it then deliberately tampers the policy to grant the operator read access and
asserts that the script catches it.
That's the whole thesis: don't ask to be trusted, ship the thing that makes trust unnecessary.
Two flows, which are the product
Strip away the infrastructure and Sepulchre is two verbs.
Inbound — "send me a secret." An operator creates an intake: a named request ("AWS production keys") with a list of fields, aimed at a recipient. The broker mints a single-use submission link carrying a scoped JWT and emails it. The recipient opens a minimal form, types the values, and submits — and the values go straight into Vault, at the operator-denied path. The operator's console shows the intake flip from "pending" to "fulfilled," and that is all they ever see. When a machine needs the credential, it presents its AppRole token to the one gated endpoint that returns inbound values, and every such read is written to the audit log.
Outbound — "here's a secret, once." An operator types a secret value into the console. The broker writes it to Vault, wraps it, and returns a reveal link. The recipient opens it, sees the secret exactly once, and the underlying secret is destroyed. A second visit — by the legitimate recipient who refreshed, or by an attacker who intercepted the link — gets a generic error and trips an audit event flagged as a possible interception. The single-use property isn't a UI convenience; it's the tripwire that tells you a link leaked.
Everything else in the system exists to make those two verbs safe.
The shape of the thing
Recipient browsers Operator browser
┌──────────┴──────────┐ │
Intake portal Reveal page Operator console
(Astro, zero-JS) (Astro, minimal JS) (Astro + islands)
└──────────┬──────────┴──────────────────────────────┘
▼ same-origin only
┌───────────────┐
│ Broker (Hono) │ authn/z · workflow · audit emitter
└───────┬───────┘
┌──────────┼───────────────┐
▼ ▼ ▼
Vault PostgreSQL Audit sinks
KV + wrap metadata only (Postgres + JSONL)
The broker is a small TypeScript service (Hono, Drizzle, a thin typed Vault client). It is the only thing that talks to Vault and Postgres. It is deliberately not clever: it has no function anywhere in its codebase that reads inbound data with its own privilege. When a service account reads a submission, the broker performs that read as the caller's token, so Vault's policy — not the broker's good intentions — is the authority. This matters for the threat model: a fully compromised broker host can mint submission links and create bogus shares, but it still cannot read inbound credentials, because it never held that capability in the first place.
The three front-ends are separate Astro applications, one per trust context, server-rendered and served from their own subdomains. Splitting them isn't aesthetic. The intake portal and reveal page are where untrusted recipients meet credentials, so they get the tightest content security policy in the system and the least code. The console is where an authenticated operator works, so it can afford interactivity. Keeping them as separate origins means a flaw in the rich console can't reach into the austere credential-handling pages.
Postgres holds metadata only. This is a hard rule, stated at the top of the schema and enforced by review and by a runtime guard: no secret values, no Vault tokens, no wrap tokens, nothing that decrypts a secret. The only secret-adjacent columns are hashes (argon2id password and passphrase hashes) and opaque accessors (a Vault wrap-token accessor, which is a handle for revocation, not the token). If the database is exfiltrated wholesale, the attacker gets a list of who-asked-whom-for-what and when, and not a single credential.
Design decisions worth defending
A system is its trade-offs. Here are the ones that shaped Sepulchre.
Response wrapping, and the accessor as a capability
For outbound shares, the obvious move is to put a Vault response-wrapping token in the URL and let the recipient unwrap it. Vault's wrap tokens are single-use by construction, which is exactly the property you want. But two requirements complicate it: shares can allow up to five views, and Postgres must never store the token (storing it would let a DB leak unwrap the secret).
The resolution is to treat the wrap token's accessor — a high-entropy, opaque handle that is useless without a Vault token to act on it — as the public capability in the reveal URL. The secret value lives in Vault; the broker enforces the view count in Postgres and deletes the underlying secret when the last allowed view is spent. The accessor is safe to expose because, in the hands of someone with no Vault access, it does nothing. The single-use security property is reproduced by broker-side view counting plus deletion, and the interception signal — an attempt after the views are exhausted — becomes an explicit, audited event. It's a pragmatic reading of Vault's primitives that keeps the "nothing in Postgres decrypts a secret" invariant intact.
The same-origin proxy, because strict CSP and CORS don't mix
The reveal page needs a little JavaScript: fetch the secret on an explicit click, show a copy
button, run an expiry countdown. But its content-security-policy is strict — connect-src 'self',
no third-party origins — and the broker lives on a different subdomain. A direct browser fetch
would be cross-origin, which means relaxing CSP and dealing with CORS, both of which widen the
attack surface on the most sensitive page in the system.
Instead, each front-end proxies the broker through its own origin. The reveal page's client
script calls /api/reveal/:accessor on its own host; a tiny Astro server endpoint forwards that
to the broker over the internal network and relays the response. The browser only ever talks to
'self'. The secret transits the Astro server, never touching its logs, and the page keeps the
tightest CSP it can. A nice side effect: the secret arrives via JavaScript and lives only in the
DOM, so a refresh reloads a secret-free shell and the second unwrap fails — the one-time property
falls out of the architecture rather than being bolted on.
Zero JavaScript on the portal, and what "strict" really costs
The intake portal goes further: script-src 'none'. No inline scripts, no external scripts, no
framework — a server-rendered HTML form and nothing else. The "submit only when valid" behavior
that normally needs JavaScript is delegated to native HTML constraint validation. This sounds easy
until you discover that Astro, trying to be helpful, inlines small stylesheets — and an inline
<style> violates style-src 'self' just as surely as an inline script violates script-src. The
fix was to force every stylesheet external. The lesson generalizes: a strict CSP is only as strict
as your build pipeline's defaults, and you have to verify the rendered bytes, not the policy you
intended.
Secrets must never reach a log — enforced three ways
The most boring way to leak a credential is to log it. Sepulchre treats "a secret reaches a logger, a response body it shouldn't, or a queue" as a class of bug to be made structurally hard, in three layers:
- A type boundary. Values read back from Vault are branded
Sensitive<T>, so every place one flows to a sink is an explicit, greppableexposeSensitive()call. The brand is advisory — TypeScript can't stopString(x)— but it makes intent legible at the boundaries. - A static gate. A semgrep ruleset, run fail-closed in CI, taints any value from a Vault read (or a parsed request body) and fails the build if it reaches a logger. It ships with a fixture of deliberate leaks, and CI asserts the ruleset still flags all of them — so the gate can't rot into a no-op.
- A runtime guard. The audit emitter rejects metadata containing known secret-bearing keys
before any sink sees the event, and the request logger emits only the templated route
(
/api/v1/reveal/:id) — never the raw path, because the path segment is itself a capability token.
None of these is sufficient alone. Together they make the easy mistake hard to commit, hard to merge, and hard to execute.
Audit as evidence, not telemetry
In most systems the audit log is an afterthought you consult during an incident. Here it's part of
the product's value proposition: the zero-knowledge claim is only as strong as your ability to show
who actually read what. Every operation emits an event with a typed actor — human, recipient,
service, system — to two sinks: Postgres (queryable) and a JSON-lines file (portable). The
service-account read of an inbound secret is audited. The second reveal attempt is audited and
flagged. For a clean tenant, the list of human reads of inbound data is empty, and provably so.
The proof you can run
It's worth restating, because it's the part that makes the rest credible. From a provisioned instance:
$ bash tools/verify-policy.sh
==> [t_default] comparing live 'operator-t_default' policy to the rendered published template
OK: live policy matches the published artifact (tenant t_default)
==> [t_default] asserting an operator-scoped token is DENIED reading inbound data
OK: operator-t_default denied on sepulchre/data/t_default/inbound/* (zero-knowledge holds)
==> zero-knowledge property verified.
Two independent checks: the live policy is the published one (so there's no rug-pull between the file you reviewed and the system that's running), and an operator token is behaviorally denied (so the property is real, not merely declared). This runs in CI on every change. A pull request that loosens the operator policy — anywhere — fails the build.
What the POC deliberately is not
Honesty about scope is part of a security tool's credibility. The POC is one tenant, two flows, runnable on a single box, with the zero-knowledge property genuinely enforced. A pile of hard, important things are explicitly deferred:
- BYOK and pluggable root-of-trust. The strongest form of the claim is a tenant supplying their own KMS seal key, so that revoking it makes their data cryptographically inaccessible including to the operator. The POC ships the policy-layer guarantee; the cryptographic one is next.
- Hash-chained, WORM-archived audit. Today's audit is durable across two sinks but not yet tamper-evident. A signed hash chain plus object-lock archival is what turns "we logged it" into "we can prove the log wasn't edited."
- HA and auto-unseal. The POC Vault uses file storage and Shamir unseal with operator-held keys. Production wants a Raft cluster and auto-unseal.
- SSO, multi-tenancy hardening, SDKs. All real, all later.
Shipping the guarantee first and the scale second is the right order for a tool whose entire pitch is trustworthiness.
What building it taught us
The most instructive moment came at the very end, driving the seven acceptance criteria against a real, running stack instead of mocks. Two bugs surfaced that every unit test had cheerfully missed.
The first was an audit sink that threw on its first write because its target directory didn't
exist — invisible in tests (which write to /dev/null) and fatal on a fresh deploy. The second was
sharper: the service-account read path didn't work in the assembled application at all. Its unit
test passed because it mounted that route in isolation. But in the real app, the operator sub-app
guarded itself with a catch-all middleware that shadowed the sibling service route, so a perfectly
valid machine token was rejected by the human-session guard before the service gate ever ran. The
flagship feature — the one endpoint that returns inbound credentials — was unreachable, and nothing
in the test suite knew.
The fix was small (scope the middleware to its own paths). The lesson was not: integration is a distinct property from unit correctness, and the only way to know your system works is to run the whole thing and watch a real secret travel from one stranger to a machine and never to the operator in between. We added a full-application regression test so that particular ghost can't return, and we kept the acceptance run in the documentation, because catching that bug is exactly what an end-to-end demo is for.
Closing
Sepulchre is a small system with one strong opinion: when a secret has to pass between people, the broker in the middle should be structurally unable to read it, and should be able to prove it. The mechanism is unglamorous — a Vault deny policy, a service that holds no read capability, metadata- only storage, and a script that checks the live system against the published rule. The unglamorousness is the point. Trust that rests on a verifiable, falsifiable artifact is worth more than trust that rests on a privacy page, precisely because you don't have to extend it.
The next chapters — BYOK, tamper-evident audit, the scale-out — are about deepening that guarantee and making it survive real production. But the foundation is in place, and it does the one thing the name implies: it keeps what you put in it, and tells no one.