Secrets Management Beyond .env Files: Vault, External Secrets, and CI Hygiene

Every codebase that touches a database, an S3 bucket, or a third-party API has secrets, and every team eventually learns the same lesson: secrets leak. They leak through git history, CI logs, Slack screenshots, .env files copied to laptops, container image layers, and kubectl describe pod output. Rotating a leaked key is painful; rotating one that is baked into forty deployment manifests across three clusters is a project.

The goal of secrets management is not to find a perfect hiding place for credentials. It is to minimize the number of places a secret exists, automate its delivery to the workloads that need it, and make revocation cheap. This post walks through the architecture that achieves that in practice: a central store, workload identity instead of long-lived credentials, GitOps-safe manifests, and hygiene rules for the places secrets still end up — CI logs above all.

The Core Architecture: Store, Identity, Delivery

Mature setups converge on three components. First, a central store that keeps secrets encrypted at rest, versions them, and logs every read. That can be HashiCorp Vault, AWS Secrets Manager, Google Secret Manager, or Azure Key Vault — the specific product matters far less than the properties: versioning, audit logs, and API-driven access. Second, workload identity: the workload proves what it is to the store using a platform-issued token rather than a static credential. Third, delivery: secrets reach pods or processes at runtime, either fetched by the application or injected by a controller.

The identity piece is where most of the security actually comes from. On AWS, EKS Pod Identity or IRSA lets a service account assume a role; on GCP, Workload Identity Federation does the same. Vault supports a Kubernetes auth method where pods exchange a short-lived projected token for Vault credentials. The win is that the only credential a pod ever holds is short-lived and scoped. If it leaks, it is already expiring; if the pod is compromised, the blast radius is one service account’s policy — not a root key in an environment variable that lives for a year.

Once you have identity, applications can fetch secrets directly from the store at startup. That is a perfectly good pattern. Its weakness is that application code now carries the fetch logic, and the secret still ends up in process memory and often in environment variables. The alternatives — injection controllers and init containers — move the logic out of the app, which is what most teams end up wanting.

Kubernetes-Native Injection: External Secrets and Sealed Secrets

Two open-source projects dominate the GitOps-friendly space, and they solve different halves of the problem.

External Secrets Operator (ESO) reconciles Kubernetes Secret objects from external stores. You define an ExternalSecret custom resource that names a store and a key; the operator fetches the value and writes a regular Secret that your pods mount. The cluster manifest contains no secret material at all — just references. Rotation becomes a store-side operation: bump the version in AWS Secrets Manager, and ESO’s refresh loop pushes it to clusters. ESO supports a long list of providers and supports templating, so one stored secret can fan out into differently-shaped Secrets per environment.

Sealed Secrets attacks the inverse problem: you have a secret value that is not in an external store (a one-off API key, a TLS keypair) and you want it in git without being readable there. The controller holds a private key; you encrypt locally with kubeseal, commit the resulting SealedSecret, and the controller decrypts it into a Secret inside the cluster. Encryption is asymmetric, so the git copy is useless to an attacker who has not compromised the controller’s key. The trade-off against ESO is rotation: rotating means re-encrypting and re-committing, which is exactly the manual toil the central-store model avoids.

The pragmatic combination: ESO for everything that lives in a real store, Sealed Secrets (or SOPS-encrypted files) for the handful of bootstrap secrets — including the credentials ESO itself needs to talk to the store. That bootstrapping corner is the classic chicken-and-egg of GitOps secrets, and it is the one place where an encrypted-in-git value is the right tool.

SOPS deserves a note even outside that niche: it encrypts values in YAML/JSON files while leaving keys readable, so diffs and reviews still work. It integrates with KMS, GCP KMS, or age keys, and its git-diff-friendly design is why many teams prefer it to opaque encrypted blobs for config that lives in git.

CI/CD: The Biggest Leak Surface Nobody Owns

Pipelines are where secrets go to die. A deploy key with write access, a cloud credential scoped to everything, a docker login whose token echoes into the log — every one of these has shipped real incidents. Four rules cover most of it.

  • Use the platform’s secret store, not variables passed between steps. GitHub Actions secrets are masked in logs, but masking is a last line of defense, not a guarantee — it will not catch base64-encoded values, substrings, or values written to files that later get printed.
  • Scope credentials to the minimum. A deploy credential needs write access to one registry namespace, not the whole cloud account. For GitHub Actions on AWS, OIDC federation (the same workload-identity idea as above) removes long-lived AWS keys from CI entirely.
  • Treat fork pull requests as untrusted. pull_request_target workflows run privileged by design — they get secrets and repository write access — so the footgun is checking out and building untrusted PR code inside that privileged context. Read the trust model of your CI’s PR events before wiring deployments to them.
  • Assume logs leak and design for revocation. Short-lived credentials in CI are not just hygiene; they are the reason an accidental echo is an annoyance instead of a week of rotation.

Detection and Hygiene

No architecture survives contact with a hundred developers, so detection is part of the design. Server-side: secret scanners on every push (the store should also alert on reads from unusual principals — an audit log nobody reads is decoration). Client-side: pre-commit scanning catches accidents before they are pushed, which matters because a secret removed in a follow-up commit is still in git history and still needs rotation. Treat any secret that ever touched a remote as burned — rotation is the only remediation, and force-pushing history away does not help once a mirror or fork exists.

Two hygiene rules round this out. First, rotate on a schedule, not just on incident: if credentials are rotated quarterly and delivery is automated, rotation is boring; if it has never happened, the first rotation will be an outage. Key management lifecycles are covered in NIST SP 800-57 if you need a formal reference for your policy. Second, kill the exceptions list: every “just this once” hardcoded credential becomes the persistent one that outlives the team that wrote it. Misconfiguration of exactly this kind — secrets in code, overly broad credentials, missing encryption — sits inside OWASP’s Security Misconfiguration category in the OWASP Top 10:2025, and it remains one of the most commonly exploited classes in practice.

Environment variables deserve an explicit mention because they are the default habit. They work, but they leak in predictable ways: child processes inherit them, crash dumps and /proc exposure reveal them, and error reporters happily serialize the whole environment. Reading from a mounted file or fetching at startup from the store is slightly more work and meaningfully harder to leak. If you must use env vars, make sure your crash reporting and log redaction are configured before the first incident, not after.

Wrapping Up

The architecture that holds up under scale is not exotic: a versioned central store, short-lived workload identity instead of static keys, runtime injection via External Secrets (with Sealed Secrets or SOPS for the bootstrap corner), tightly scoped OIDC-federated CI credentials, and scanners on both sides of the push. Each piece is boring on its own; together they make secret rotation a non-event and make leaks small instead of catastrophic. If your current setup is a .env file and a Slack DM, start with the store and workload identity — everything else gets easier once those two exist.

Leave a Reply

Your email address will not be published. Required fields are marked *