etcd Encryption
Install and operate etcd Encryption Manager in to automate etcd data encryption key rotation within your clusters.
It ensures that sensitive data stored in etcd, such as secrets and configmaps, is encrypted using a secure algorithm, enhancing your cluster's security.
TOC
Key StrategiesInstallationWorkload / DCS Clusters (random strategy)Global Clusters (Deterministic Mode)PrerequisitesPlugin ParametersPreparing the Root SecretAdvanced (chart values, not exposed in the plugin form)How it WorksDefault ConfigurationOperations GuideConfiguration FilesChecking StatusApproving a Deterministic Key RotationManual Approval (Recommended)Timed Approval (Not Recommended)Key Strategies
The plugin supports two key strategies. Pick the one that matches your cluster role.
In deterministic mode, the runtime Active / Standby role is detected automatically; it is not a manual install parameter.
Installation
The plugin supports two installation paths depending on the target cluster:
- Workload / DCS clusters — use the default
randomkey strategy. - Global clusters — must use the
deterministickey strategy (see Global Clusters (Deterministic Mode) below).
See Cluster Plugin for the general plugin installation workflow.
Workload / DCS Clusters (random strategy)
- Supported cluster types: On-Premises, DCS.
- No additional installation parameters are required — install via the standard Cluster Plugin workflow.
Global Clusters (Deterministic Mode)
For a Global DR pair, install the plugin on both the Active and Standby clusters in deterministic mode, and configure identical replication_group_id and root_secret_name (with matching root key material) on both sides. Asymmetric or mismatched configuration causes the Standby to derive different per-revision keys than the Active and breaks failover.
Prerequisites
- The cluster must be a
globalcluster. Whenglobal.cluster.nameisglobal, the plugin enforcesetcdEncryption.keyStrategy=deterministic; the plugin configuration UI also defaults the field todeterministicfor clusters labeledis-global=true.
The etcd Synchronizer plugin must be installed at version v4.3.7 or later before installing the etcd Encryption Manager in deterministic mode. Earlier versions do not replicate the shared SeedBundle, so the Standby cluster cannot derive matching keys.
Plugin Parameters
The plugin form exposes the following parameters for deterministic installation.
The rootSecretRef.namespace is fixed to kube-system and is not exposed in the form.
key_strategy
For global clusters this must be deterministic.
replication_group_id
Recommendations:
- For DR clusters built with the etcd Synchronizer, both sides of the same DR group must share the same
replication_group_id. - To avoid leaking real business, environment, or cluster identity, do not reuse the cluster name even when the UI prefills it. Plan a separate opaque identifier ahead of installation.
- A UUID-style identifier may be a good fit, for example
6f1b9e2c-7c3a-4a9c-8b72-2fd0f8f3c1ab.
To generate a UUID-style value from the command line:
Generate the value once, then configure the same replication_group_id on the Active and Standby clusters of the DR group. Generating different values on each side will produce divergent encryption keys.
root_secret_name
The plugin form exposes only the Secret name; the namespace is always kube-system at render time.
Preparing the Root Secret
The deterministic root Secret must exist in kube-system before installing the plugin. Prepare it in two steps:
-
Generate 32 bytes of cryptographically strong random material into
./root-key.bin. The bytes must come from a high-entropy source — typically the operating system CSPRNG.openssl randis a safe default; equivalent OS-CSPRNG-backed commands such ashead -c 32 /dev/urandom > ./root-key.binalso work. Do not use shell$RANDOM, language-levelrand(), or any time-seeded generator. -
Create the Secret in
kube-systemfrom that file:
If you use a different Secret name, replace etcd-derivation-root above with the value you will enter as root_secret_name.
Generate ./root-key.bin once, then transport the same file to every cluster in the replication group and run the kubectl create secret step on each. Do not regenerate the key material per cluster — different root key bytes will produce divergent encryption keys.
Advanced (chart values, not exposed in the plugin form)
Show advanced chart values
The following fields exist in the chart but are not surfaced as plugin install parameters. They retain their chart defaults unless you install via YAML / Helm overrides.
For a Global Active/Standby pair, explicitly set activationPolicy to manual. A new key revision must not become active on the Active cluster until an operator has confirmed that the matching revision is ready on the Standby cluster. If activationPolicy is omitted, the plugin defaults to timed-approval with a 30m delay — see Timed Approval for the risks this carries.
How it Works
Upon installation, an etcd-encryption-manager controller is deployed in the kube-system namespace, which:
- Periodically rotates etcd data encryption keys.
- Retains the 8 most recent keys for rollback compatibility.
- Updates encryption configurations on all control nodes.
- Triggers
kube-apiserverto hot reload new keys. - Automatically migrates resources to re-encrypt data with new keys.
Cluster stability is maintained throughout these operations.
When the plugin is installed in deterministic mode on a global DR pair, the controller additionally:
- Detects whether it is acting as Active or Standby at runtime — this is not a manual configuration.
- On the Active side, generates the SeedBundle and lets the etcd Synchronizer replicate it to the Standby.
- On the Standby side, derives the same per-revision keys from the shared SeedBundle and the root Secret, so a failover produces identical encryption keys without manual key copy.
Default Configuration
Operations Guide
Configuration Files
Checking Status
Run the following command to check the current rotation status:
Example output:
Approving a Deterministic Key Rotation
When the encryption manager rotates a key in deterministic mode, the new revision enters a WaitingApproval phase on the Active cluster. The key does not take effect until the operator approves it. This gate exists because the Active cluster must not encrypt data with a key that the Standby cluster has not yet loaded — otherwise a failover would leave Standby unable to decrypt the new ciphertext.
Approval is always performed on the Active cluster only. Do not patch the Standby's EtcdKeySeedBundleState, and do not edit any resource's status field directly.
Manual Approval (Recommended)
Step 1 — Identify the pending revision.
Log in to a control-plane node of the Active cluster and list all bundle states:
Find the row with PHASE=WaitingApproval and APPROVED=false. Record its revision number:
Step 2 — Verify Standby readiness.
Log in to a control-plane node of the Standby cluster and check the same revision:
The output must show PHASE=Prepared for the expected revision. Stop if the object is missing, still Preparing, or Failed — investigate the etcd Synchronizer path before proceeding.
Step 3 — Cross-check derivation inputs.
Run the following command on both the Active and Standby clusters and compare the output. The replicationGroupID, dimension, and derivationAlgorithm must match exactly:
Step 4 — Approve.
On the Active cluster, approve the pending revision:
The patch records the approval. Activation, deployment to control-plane nodes, and data migration happen asynchronously after this point.
Step 5 — Confirm activation completes.
On the Active cluster, check the bundle state and migration progress:
The rotation is complete when:
- The bundle state reaches
PHASE=Active. eec/defaultshows the newREVISIONwithMIGRATION_STATE=Success.
To check per-node deployment progress on the Active cluster:
Every control-plane node must report state=Success for the new revision. A bundle that remains Activating indicates deployment or migration is still in progress (or retrying after a failure).
Timed Approval (Not Recommended)
timed-approval is not recommended for Global Active/Standby pairs. The approvalDelay timer (default 30m) is a local countdown on the Active cluster — it does not check whether the Standby has actually loaded the new key. If etcd synchronization is slow, the Standby network is partitioned, or the Standby control plane is under maintenance, the deadline can expire while the Standby is still unprepared. A failover during this window leaves Standby unable to decrypt data encrypted with the new key.
When activationPolicy is set to timed-approval, the controller automatically approves the pending revision once approvalDelay elapses after the EtcdKeySeedBundleState is created. The delay must be long enough for the full Standby preparation pipeline to complete:
The default 30m provides margin for typical single-region deployments with a small control plane. Increase it if your setup involves cross-region replication, large control planes, or clusters without automatic provider reload.
Required monitoring
If you choose timed-approval, you must verify that monitoring and alerting are working on both the Active and Standby clusters before enabling this policy.
The encryption manager exposes the following Prometheus metrics:
Metrics
Built-in alert rules
The plugin includes the following alerts:
These alerts fire on each cluster independently. For Global clusters, the critical alert to act on is EtcdEncryptionManagerTimedApprovalDeadlineApproaching on the Active cluster — it gives you a 10-minute window to check whether the Standby is Prepared and switch to manual if it is not.
Each cluster runs its own metrics and alerting independently. The built-in alerts detect local issues — a stuck WaitingApproval on Active or a failed activation on Standby. To correlate the two (e.g., "Active is about to auto-approve but Standby is not yet Prepared"), aggregate metrics from both clusters into a central Prometheus / Thanos / Victoria Metrics instance.
Emergency response
If the EtcdEncryptionManagerTimedApprovalDeadlineApproaching alert fires and the Standby is not ready, switch the Active to manual to halt auto-approval. On the Active cluster:
Verify the current pending bundle has switched:
The bundle must still show APPROVED=false. Switching back to timed-approval later affects only future bundles — the current bundle remains under manual control. If the bundle has already moved to Activating, the deadline has already taken effect; follow the activation checks above and do not assume that changing the policy rolls it back.