Skip to content

stack-encrypt: one context per column — render descriptors with / and bind (table, column) as a pair everywhere #1049

Description

@coderdan

Background

stack-encrypt (packages/stack-encrypt) encrypts each value under its own data key from ZeroKMS, our key management service. Every value is encrypted under a context: a small structured value that says what the data is, such as a table and column. A context has parts (text, bytes, integers, or lists of those), and stack-encrypt turns it into two different encodings:

  • The AAD and PRF context (what the ciphertext is bound to, and what separates search terms per column) use vitaminc's PAE encoding. It's binary, records each part's length and type, and can never be ambiguous.
  • The ZeroKMS descriptor is a readable string rendered from the same parts (Descriptor::from_piece, packages/stack-encrypt/src/descriptor.rs). ZeroKMS takes a string, binds it into the data key's tag, and writes it to the retrieval log. That log is how an audit trail shows which column was read.

The descriptor joins a context's parts with Descriptor::SEPARATOR, which is | (descriptor.rs:71). A text part containing |, (, ) or a control character is base64-escaped, which keeps the rendering one-to-one.

Problem

The same column is spelled two ways, so it renders two different descriptors and two different contexts.

  • The #[derive(EncryptFrom)] macro binds a field under the one-part text "users/email", joined with format!("{prefix}/{column}") (packages/stack-encrypt-derive/src/shape.rs:369). Its descriptor is users/email.
  • The EQL v3 bindings in feat(eql): add Rust text encryption and queries #971 bind a column's Identifier as the two-part pair ("users", "email"). Its descriptor is users|email.
  • The Go binding mirrors the derive. Its plan contexts are joined strings (FieldPlan.Context string, languages/golang/stackencrypt/record.go:207), built by plan.Identifier.String() as table + "/" + column (plan/policy.go:26).

The two spellings are different contexts in every encoding. The AAD differs, so neither opens the other's ciphertext. The PRF differs, so search terms never match. And ZeroKMS refuses to release a key minted under users/email to a request for users|email. A Rust EQL reader and a Go or derive writer of the same column wouldn't interoperate, and nothing fails until data is read or searched.

Joining parts into one string also pushes a rule onto every caller: table and column names may not contain /, or ("a/b", "c") and ("a", "b/c") become the same string. The Go plan already enforces that by hand (plan/message.go).

Nothing has been published yet (stack-encrypt goes to crates.io at 0.1.0 in #1047), so the rendering can still change without stranding any keys.

Proposal

  1. Make / the descriptor separator, and escape it inside text. Change Descriptor::SEPARATOR to '/' and add / to the characters that force a text part into its base64 form, so the rendering stays one-to-one. The multi-part pair ("users", "email") renders users/email.
  2. Make contexts structured, not pre-joined:
    • The derive: struct = User, context = "users" binds each field under the pair ("users", "<column>").
    • eql-bindings' Identifier (feat(eql): add Rust text encryption and queries #971): binds the pair (table, column), which it already does.
    • The Go binding: a plan carries each field's context as a structured value, not a joined string, and the struct-tag plan mirrors the derive's pair. Go's Context already spells a pair (NewContext("users").With("email")).
  3. Update the frozen-rendering tests and every pinned descriptor and context in stack-encrypt, the derive and Go, including the cross-language vectors Go checks against the Rust derive.
  4. Land it before stack-encrypt 0.1.0 is published, i.e. before chore(release): publish stack-kms, stack-encrypt and stack-encrypt-derive to crates.io #1047's hand publish, so the first published version carries the final descriptor format.

After this, every column renders the same descriptor everywhere, users/email. That matches cipherstash-client's existing table/column descriptors, and table and column names no longer need a "no /" rule.

Relationship to other work

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

SDKenhancementNew feature or requestrustPull requests that update Rust code

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions