Vault Healthcheck

The Vault endpoints that report seal status, HA state, Raft health and DR replication, and how to read them before maintenance.

Checking whether Vault is running is easy. Knowing whether it is safe to touch the infrastructure around it is not.

/sys/health answers one question: what state is this node in? It says nothing about how much redundancy the Raft cluster has left, whether the HA nodes agree on who is active, how far a DR secondary has fallen behind, or whether anyone has taken a snapshot recently. Each of those lives in a different endpoint, and any of them can be the reason a maintenance window goes wrong.

This note collects them: what each endpoint reports, what the fields in its response mean, and a short script that collects evidence for review before updating a Kubernetes cluster, draining a node, or starting VMware maintenance.

One caveat up front: the health and seal endpoints apply to any Vault, but the Raft, Autopilot, and snapshot commands assume Integrated Storage. On a Consul or other storage backend, those sections do not apply.

Contents↗

Why Health Checks Matter↗

A health check that only distinguishes “up” from “down” is not enough for Vault. Vault has states that are neither: sealed but running, unsealed but standby, active but disconnected from its replication primary. Each one calls for a different response, and a binary check collapses them into noise.

The second reason is routing. In an HA configuration, a load balancer needs to send traffic to the active node and keep standby nodes in rotation as candidates, not as failures. Vault is designed for this: it returns distinct status codes per state so the balancer can tell “up but not the one you want” apart from “down”.

The third is timing. Clock drift between nodes and replication lag degrade before they break. They are visible in these endpoints while they are still cheap to fix.

Vault is rarely either up or down. The useful question is which state it is in, and whether that state is the one you expected.

Prerequisites and Scope↗

To collect all the evidence and save a snapshot, the example needs:

  • Vault configured with Integrated Storage;
  • the Vault CLI on the operator host, plus curl for the status-code examples;
  • a Vault token with the capabilities required for Raft, Autopilot, HA status, DR status where available, and snapshot creation;
  • an existing, writable /opt/vault directory on the operator host;
  • to be run from the root namespace, against the cluster you intend to maintain.

DR replication is Enterprise-only. A failed command remains visible in the output while the script continues collecting the other evidence.

The scope is control-plane state on the cluster you address: seal status, HA role, Raft topology, replication, and a snapshot. What it leaves out is covered at the end.

The /sys/health Endpoint↗

/sys/health is the primary endpoint for operational status. It reports whether Vault is initialized, unsealed, active, standby, or replicating — and it encodes the answer in the HTTP status code, so monitoring tools can act on the response without parsing the body.

Status Codes↗

CodeStateWhat it means
200Initialized, unsealed, activeReady to process requests allowed by authentication and policy
429Unsealed, standbyFunctional but not the active node. The odd choice of code is deliberate: it keeps standby nodes out of a balancer’s active pool without marking them unhealthy
472DR replication secondaryA healthy state, not a failure. A DR secondary is supposed to exist and to refuse normal traffic
473Performance standbyEnterprise. Serves read-only requests locally and forwards writes to the active node
474Standby, no connection to the active nodeThe node is not a usable failover candidate, even if the cluster still serves traffic. Worth alerting on
501Not initializedExpected on a new instance until vault operator init runs. Later, it usually means the request reached the wrong instance
503SealedNo secret can be read or written until Vault is unsealed. The critical condition for most monitoring
530Removed from the HA clusterThe node is no longer a cluster member

Two of those are routinely misread. 429 and 472 are healthy states: a standby node and a DR secondary are both working as designed. Monitoring that treats them as outages pages on a correctly configured cluster.

If the defaults do not suit your load balancer, the codes are remappable through query parameters: activecode, standbycode, sealedcode, drsecondarycode, performancestandbycode, haunhealthycode, uninitcode, and removedcode. The standbyok and perfstandbyok flags can also map those healthy standby states to the active code. Setting standbycode=200, for example, makes standby nodes indistinguishable from active ones to a balancer that only understands 2xx.

Accessing the /sys/health Endpoint↗

The endpoint is part of Vault’s HTTP API and requires no authentication, which makes it usable from load balancers and probes that hold no token.

Because the state lives in the status code, a plain curl is not enough — it prints only the body. Ask for the code explicitly:

# Status code only, suitable for scripting
curl -s -o /dev/null -w '%{http_code}\n' "$VAULT_ADDR/v1/sys/health"

# Code and body together
curl -i "$VAULT_ADDR/v1/sys/health"

The CLI can display fields from the same endpoint, but it does not expose the raw HTTP status as directly and may report non-2xx states as request errors. Use curl when the status code is the value being tested:

vault read sys/health

Example JSON response:

{
  "initialized": true,
  "sealed": false,
  "standby": false,
  "performance_standby": false,
  "replication_dr_mode": "disabled",
  "replication_performance_mode": "disabled",
  "server_time_utc": 1730844191,
  "version": "1.14.4",
  "cluster_name": "vault-cluster-7cb803e8",
  "cluster_id": "c792cdf2-2ff3-7ca1-0835-31667b14a9b5"
}

Example CLI response:

Key                             Value
---                             -----
cluster_id                      c792cdf2-2ff3-7ca1-0835-31667b14a9b5
cluster_name                    vault-cluster-7cb803e8
initialized                     true
performance_standby             false
replication_dr_mode             disabled
replication_performance_mode    disabled
sealed                          false
server_time_utc                 1730844191
standby                         false
version                         1.14.4

Note that standby and performance_standby are separate fields, and that replication_dr_mode tells you whether the 472 code is even reachable on this cluster.

Seal Status↗

/sys/seal-status answers the question /sys/health compresses into a single code: if Vault is sealed, how far has the unseal progressed. Vault data remains encrypted at rest in either state. When sealed, Vault cannot recover the key material needed to decrypt that data, so no secret is accessible until enough key shares have been supplied or the auto-unseal mechanism succeeds.

vault read -format=json sys/seal-status

Response example:

{
  "request_id": "",
  "lease_id": "",
  "lease_duration": 0,
  "renewable": false,
  "data": {
    "build_date": "2023-09-22T21:29:05Z",
    "cluster_id": "c792cdf2-2ff3-7ca1-0835-31667b14a9b5",
    "cluster_name": "vault-cluster-7cb803e8",
    "initialized": true,
    "migration": false,
    "n": 1,
    "nonce": "",
    "progress": 0,
    "recovery_seal": false,
    "sealed": false,
    "storage_type": "inmem",
    "t": 1,
    "type": "shamir",
    "version": "1.14.4"
  },
  "warnings": null
}
FieldMeaning
sealedtrue means Vault is locked and inaccessible; false means it is unsealed and serving
tThreshold: the number of key shares required to unseal. Set at initialization
nTotal number of key shares issued, typically distributed across several holders
progressShares entered so far in the current unseal attempt. Vault stays sealed while progress < t, and unseals when it reaches t
nonceIdentifies the current unseal attempt. Shares from different nonces do not combine
typeThe seal mechanism: shamir for key shares, or awskms, azurekeyvault, gcpckms and others for auto-unseal
recovery_sealtrue when the instance uses auto-unseal and is reporting recovery-key state rather than unseal-key state
migrationtrue while a seal migration is in progress, for example moving from Shamir to auto-unseal
storage_typeThe configured storage backend, such as raft, consul, or inmem
versionVault version, useful when behavior differs across releases
cluster_name, cluster_idCluster identifiers, useful in multi-cluster and DR setups

type matters operationally more than it looks. With shamir, unsealing needs t humans and their key shares. With an auto-unseal backend, it needs the KMS to be reachable — which makes the KMS a dependency of your secret store, and a sealed Vault a possible symptom of a cloud IAM problem rather than a Vault one.

Raft Autopilot State↗

/sys/storage/raft/autopilot/state reports the state of the integrated storage cluster: each node, its role, and how much redundancy remains. Autopilot manages voter stabilization and promotion. It can also remove failed servers when dead-server cleanup is explicitly enabled. These controls reduce manual peer management, but they do not make quorum self-healing under every failure.

vault operator raft autopilot state -format=json

Response example:

{
  "healthy": true,
  "failure_tolerance": 1,
  "servers": {
    "raft1": {
      "id": "raft1",
      "name": "raft1",
      "address": "127.0.0.1:8201",
      "node_status": "alive",
      "last_contact": "0s",
      "last_term": 3,
      "last_index": 459,
      "healthy": true,
      "stable_since": "2021-03-19T20:14:11.831678-04:00",
      "status": "leader",
      "meta": null
    },
    "raft2": {
      "id": "raft2",
      "name": "raft2",
      "address": "127.0.0.2:8201",
      "node_status": "alive",
      "last_contact": "516.49595ms",
      "last_term": 3,
      "last_index": 459,
      "healthy": true,
      "stable_since": "2021-03-19T20:14:19.831931-04:00",
      "status": "voter",
      "meta": null
    },
    "raft3": {
      "id": "raft3",
      "name": "raft3",
      "address": "127.0.0.3:8201",
      "node_status": "alive",
      "last_contact": "196.706591ms",
      "last_term": 3,
      "last_index": 459,
      "healthy": true,
      "stable_since": "2021-03-19T20:14:25.83565-04:00",
      "status": "voter",
      "meta": null
    }
  },
  "leader": "raft1",
  "voters": ["raft1", "raft2", "raft3"],
  "non_voters": null
}

Top-level fields:

FieldMeaning
healthyWhether every node is healthy and able to participate in quorum
failure_toleranceHow many nodes can fail while the cluster still holds quorum
leaderThe name of the current Raft leader — a string, not a per-node flag
votersNodes that participate in quorum decisions
non_votersNodes currently replicating without voting. New nodes may appear here while stabilizing; permanent non-voters for read scaling are Enterprise

Per-node fields inside servers:

FieldMeaning
id, nameNode identifier within the Raft cluster
addressCluster address of the node, on the cluster port (8201 by default)
node_statusalive, failed, or left
statusThe node’s Raft role: leader, voter, or non-voter
last_contactTime since the leader last heard from this node. On the leader itself this is 0s
last_termRaft term the node last observed. A node lagging in term is behind on elections
last_indexLast applied log index. A gap against the leader’s index is replication lag
healthyWhether autopilot considers this node healthy
stable_sinceWhen the node last entered its current state. Recent values on an old cluster suggest flapping

Gate maintenance on both healthy and failure_tolerance. The first describes whether Autopilot considers every node healthy. The second describes how many additional healthy nodes the current topology can lose without losing quorum. A cluster reporting healthy: true and failure_tolerance: 0 may not have lost a node, but it has no remaining failure budget and should not lose another one during maintenance.

healthy: true describes the present. failure_tolerance describes what the cluster can still survive.

Autopilot itself is available in all editions. Its redundancy zones and automated upgrade features are Enterprise.

HA Status↗

/sys/ha-status lists every node in the HA cluster and identifies the active one. Where /sys/health reports the state of the node you asked, ha-status gives you the cluster’s view: useful both for routing and for detecting a node that unexpectedly lost its active role.

vault read -format=json sys/ha-status

Response example:

{
  "Nodes": [
    {
      "active_node": true,
      "api_address": "http://10.0.0.2:8200",
      "clock_skew_ms": 0,
      "cluster_address": "https://10.0.0.2:8201",
      "echo_duration_ms": 0,
      "hostname": "node1",
      "last_echo": null,
      "version": "1.17.0"
    },
    {
      "active_node": false,
      "api_address": "http://10.0.0.3:8200",
      "clock_skew_ms": 0,
      "cluster_address": "https://10.0.0.3:8201",
      "echo_duration_ms": 20,
      "hostname": "node2",
      "last_echo": "2024-03-04T08:05:48.403148-05:00",
      "version": "1.17.0"
    },
    {
      "active_node": false,
      "api_address": "http://10.0.0.4:8200",
      "clock_skew_ms": -1,
      "cluster_address": "https://10.0.0.4:8201",
      "echo_duration_ms": 17,
      "hostname": "node3",
      "last_echo": "2024-03-04T08:05:48.657318-05:00",
      "version": "1.17.0"
    }
  ]
}
FieldMeaning
active_nodeWhether this node currently holds the active role. Exactly one node should report true
api_addressHTTP API address used by clients
cluster_addressAddress used for intra-cluster traffic: request forwarding and replication
hostnameNode hostname, for correlating with logs and metrics
clock_skew_msDifference between this node’s clock and the active node’s
echo_duration_msRound-trip time of the last echo check, a latency signal between nodes
last_echoTimestamp of the last echo from the active node. null on the active node itself
versionVault version on this node, useful for spotting drift mid-upgrade

On Enterprise, the response also carries replication_primary_canary_age_ms (replication lag against the primary), upgrade_version (target version during an automated upgrade), and redundancy_zone (the node’s zone in a zone-aware HA setup).

A cluster view reporting anything other than exactly one active node is the finding. It indicates an unavailable or inconsistent HA view that must be investigated before maintenance. clock_skew_ms is also worth watching: drift can affect lease timing and make health and replication observations harder to interpret.

DR Replication Status↗

Enterprise only. Disaster recovery replication requires a Vault Enterprise license. On Community edition the read returns an error, which is itself a valid observation — the script below leaves it in the output rather than hiding it.

/sys/replication/dr/status reports the state of the replication link, the connected secondaries, and how far they lag behind the primary. Checking it before maintenance confirms the DR cluster can actually take over.

The endpoint answers from both sides of the relationship: on a primary it describes the secondaries, and on a secondary it describes its own sync state against the primary.

vault read -format=json sys/replication/dr/status

Response example, queried on a primary:

{
  "data": {
    "cluster_id": "eef2a5ab-51e2-1c05-407c-8b4dc8d09ebf",
    "corrupted_merkle_tree": false,
    "known_secondaries": [
      "4ca6b639-046b-5bb1-8043-6788ddf09121"
    ],
    "last_corruption_check_epoch": "-62135596800",
    "last_dr_wal": 223,
    "last_reindex_epoch": "0",
    "last_wal": 223,
    "merkle_root": "2494830f1a1c304829b5742a232d39b5457bce9a",
    "mode": "primary",
    "primary_cluster_addr": "",
    "secondaries": [
      {
        "api_address": "https://127.0.0.1:65531",
        "clock_skew_ms": "0",
        "cluster_address": "https://127.0.0.1:65534",
        "connection_status": "connected",
        "last_heartbeat": "2024-03-04T10:05:56-05:00",
        "last_heartbeat_duration_ms": "0",
        "node_id": "4ca6b639-046b-5bb1-8043-6788ddf09121",
        "replication_primary_canary_age_ms": "696"
      }
    ],
    "ssct_generation_counter": 0,
    "state": "running"
  }
}
FieldMeaning
modeThis cluster’s replication role: primary, secondary, or disabled
stateState of the replication process. running is steady state; merkle-diff, merkle-sync and stream-wals indicate work in progress
last_walMost recent Write-Ahead Log index generated on the primary
last_dr_walMost recent WAL index replicated for DR. The gap against last_wal is the lag
known_secondariesIDs of the secondary clusters registered with this primary
merkle_rootRoot hash of the Merkle tree, a fingerprint of the replicated data state
corrupted_merkle_treetrue indicates a data integrity problem, not a lag problem
primary_cluster_addrAddress of the primary. Empty on the primary itself, populated on a secondary

Per secondary, inside secondaries:

FieldMeaning
node_idIdentifier of the secondary cluster
api_address, cluster_addressAddresses used to reach the secondary
connection_statusconnected or otherwise. The first thing to check on a failover drill
last_heartbeatTimestamp of the last heartbeat received from the secondary
last_heartbeat_duration_msRound-trip time of that heartbeat
clock_skew_msClock difference between primary and this secondary
replication_primary_canary_age_msAge of the last canary value observed, a direct lag measure in milliseconds

Connectivity alone does not clear a DR setup before maintenance. Four things have to hold: state is running, corrupted_merkle_tree is false, every expected secondary appears as connected, and the last_wal / last_dr_wal gap stays below a threshold chosen for the system’s recovery-point objective. A connected secondary can still be too far behind to satisfy that objective.

Snapshot Creation↗

vault operator raft snapshot save captures a point-in-time copy of the integrated storage: every secret, policy, mount, and piece of cluster state, written as a compressed archive. Taken before maintenance, it is a recovery artifact if something goes wrong. Restoring it is a separate, disruptive procedure, not an automatic rollback.

vault operator raft snapshot save \
  "/opt/vault/snapshot-$(date -u +%Y%m%dT%H%M%SZ).snap"

A UTC timestamp in YYYYMMDDTHHMMSSZ form sorts chronologically in any listing; %m-%d-%Y does not, which becomes annoying the first time you need the most recent snapshot in a hurry.

Snapshots are valuable for three distinct reasons:

  • Disaster recovery. Stored offsite, a snapshot is a recovery path that does not depend on the cluster that produced it.
  • Point-in-time recovery. A retained and verified snapshot provides a known recovery point when corruption or operator error appears later.
  • Recovery testing. Restoring into an isolated recovery environment verifies that the procedure, key custody, and artifact are usable. A production snapshot must not be treated as an ordinary staging fixture.

A snapshot contains the cluster’s encrypted data and still inherits the sensitivity of Vault itself. Store it outside the cluster’s failure domain, restrict access, record a checksum, define retention, and test restoration with the required unseal or recovery-key holders.

Putting It All Together↗

This script gathers evidence for a human to analyze. It prints the available operational state and attempts to save a snapshot. If a command fails, the remaining commands still run, so one unavailable endpoint does not prevent collecting the other observations.

Replace the address and token placeholders before running it. The example disables TLS verification for local development, as its comment states.

#!/bin/bash

# Disable TLS verification for local development (Not recommended for production)
export VAULT_SKIP_VERIFY=true

# Set the Vault server address (replace 'your_vault_ip' with the actual IP address of your Vault server)
export VAULT_ADDR=https://your_vault_ip:8200

# Set your Vault token for authentication (replace 'your_token' with your actual token)
export VAULT_TOKEN=your_token

# Display the list of Raft peers to verify cluster membership and health
vault operator raft list-peers

# Print general status of the Vault server, including seal status and HA state
echo "### General status"
vault status
echo ""

# Display the Raft autopilot state to check cluster health and node roles
echo "### Raft autopilot state"
vault operator raft autopilot state
echo ""

# List all peers again to ensure accurate cluster information after checking autopilot state
echo "### List Peers"
vault operator raft list-peers
echo ""

# Display Disaster Recovery (DR) replication status to confirm synchronization with DR cluster
echo "### DR status"
vault read -format=json sys/replication/dr/status
echo ""

# Display High Availability (HA) status to check the active and standby nodes in the cluster
echo "### HA status"
vault read -format=json sys/ha-status
echo ""

# Save a Raft snapshot for backup, using a sortable UTC timestamp
echo "### Snapshot"
vault operator raft snapshot save \
  "/opt/vault/backup-$(date -u +%Y%m%dT%H%M%SZ).snap"
echo ""

Script Breakdown↗

Environment. The script sets the server address, token, and TLS option at the beginning. The placeholders must be replaced for the target environment.

Command failures. There is no set -e or automatic health verdict. Review both command output and errors. The final exit status does not summarize the success of earlier commands.

General status. vault status reports seal state, storage type, version, and HA information for the addressed node.

Raft peers and Autopilot. The first peer listing records membership before the other observations. The second provides another observation later in the run. Compare both with the expected topology and review Autopilot’s member health, voter count, contact times, log indexes, and failure tolerance.

DR status. The script always attempts this read. On Community edition or when the endpoint is unavailable, an error becomes part of the evidence and the remaining commands continue. On Enterprise, interpret replication role, state, connectivity, and synchronization in context.

HA status. Review the active node and recent peer contact. The endpoint lists the active node and peers it has heard from since becoming active. Compare that view with the expected membership.

Snapshot. The CLI attempts to save the file under /opt/vault on the operator host. That directory must exist and be writable. The filename carries a sortable UTC timestamp, so repeated runs never collide and the most recent snapshot is the last one listed. Confirm snapshot creation from the command output. Retention, secure storage, and restore testing belong to the backup procedure.

What This Does Not Cover↗

Storage capacity, telemetry, and audit device health are separate concerns and are not checked here. An audit device that cannot write will block requests while every endpoint above still reports a healthy cluster.

Nor does any of this confirm that applications can authenticate. A healthy, unsealed, quorate Vault with correct replication can still be denying every request your workloads make, because auth method configuration, policies, and token lifecycles are a different layer. These observations assess Vault’s operational state; application correctness requires separate checks.

A status report is evidence. The maintenance decision still needs judgment.


More like this