Mutual TLS is easy to set up and hard to keep running, and the difficulty is almost entirely organizational. Each side holds half of the handshake. A rotation is therefore a coordinated change between two companies with different release calendars, different change-approval processes, and no shared incident channel until the thing is already broken.
Why the boundary is the problem
In a one-sided TLS setup you control your own certificate. Renew it, deploy it, done.
With mutual TLS there are four moving parts, and you own two:
| Artifact | Owned by | Breaks when |
|---|---|---|
| Your certificate | You | You renew and they have not trusted the new chain |
| Your trust of their chain | You | They renew under a chain you do not trust |
| Their certificate | Them | They renew without telling you |
| Their trust of your chain | Them | You renew without telling them |
Every one of these is a failed handshake, and a failed handshake is total: no partial service, no degraded mode, no retry that helps. The connection simply does not establish.
Separate the trust store from the key store
If you take one thing structurally, take this.
- Key store: your private key and certificate. Changes when you rotate.
- Trust store: the CA chains you accept. Changes when they rotate.
These are two different events, driven by two different organizations, on two different schedules. Putting them in one file couples them, which means every change to either requires touching the artifact both depend on — and the temptation to "just update the file" while you are in there is how the other side's rotation gets accidentally reverted.
partner:
key-store:
path: /run/secrets/client-identity.p12 # ours; rotates on our schedule
password: ${CLIENT_KEYSTORE_PASSWORD}
trust-store:
path: /run/secrets/partner-trust.p12 # theirs; rotates on theirs
password: ${PARTNER_TRUSTSTORE_PASSWORD}Trust the chain, not the leaf
Pinning the partner's leaf certificate feels more secure and turns every one of their renewals into an outage on your side.
Trust the issuing CA chain instead. Then their routine renewal under the same CA requires nothing from you at all, and coordination is only needed for the far rarer event of a CA change. If your security posture genuinely requires leaf pinning, then that requirement comes with an obligation to build the automation and the notification path that makes it survivable, and you should price that in honestly rather than discover it later.
The overlap window
The safe rotation is not a swap. It is an interval during which both certificates are valid and both chains are trusted.
┌──────────── old certificate valid ────────────┐
┌──────────── new certificate valid ────────────┐
│◀── overlap ──▶│
│ │
new chain old chain
trusted by removed
both sides by both sidesThe sequence that works, in order:
- Both sides add the new issuing chain to their trust stores. Nothing is presented yet, nothing changes behaviour, and this step is safe to do weeks early.
- Confirm. Not "we deployed it" — an actual handshake test, or at minimum the other side confirming the chain is live in the environment you will hit.
- One side switches the certificate it presents. One direction at a time.
- Verify traffic, at both the connection level and the business level.
- The other side switches.
- After the old certificate expires, both sides remove the old chain.
Step 1 is where the overlap comes from and it is the step that gets compressed when someone notices an expiry date three days out. Compressing it is how a rotation becomes an incident.
Make the window weeks, not days. It costs nothing to trust an additional chain early, and the length of the window is what absorbs the other company's change freeze that nobody told you about.
Alert on what is presented, not on what is deployed
The alert that matters reads the certificate from the live connection, not from the file you believe is in the container.
Those two diverge more often than they should: a config map that was updated but the pod never restarted, a secret mounted from the wrong path, a rollback that quietly restored an old identity. In every one of those cases the file on disk says one thing and the wire says another, and the wire is what the partner sees.
# Days remaining on the certificate the partner actually presents
echo | openssl s_client -connect partner.example.com:443 \
-cert client.pem -key client.key 2>/dev/null \
| openssl x509 -noout -enddateRun it as a scheduled check against every mTLS endpoint. Alert at 60 days, escalate at 30, page at 14. Sixty days sounds absurdly early until you have tried to get a certificate through two companies' change processes in December.
Alert on both directions where you can: your certificate as they see it, and theirs as you see it. The second one catches their forgotten renewal, and being the party who warns them buys a lot of goodwill for the day you need something.
Write down who to call
The technical runbook is the easy half. The half that actually shortens an incident is a document naming:
- who at the partner owns the certificate, with a real name and a real address
- what their change window is, and their freeze periods
- how long their approval process takes, measured rather than assumed
- where the CA chains are published
- what the expiry dates currently are, on both sides
Every mTLS incident I have watched was made longer by not having that page. The handshake failure is diagnosed in minutes. Finding a human at the other company who can act on it takes hours, and it always seems to happen on a Friday.
Related
The transport layer is only one of three in a typical enterprise integration — ingesting HSM-encrypted files over SOAP with mTLS and WS-Security covers how it fits with message signing and payload encryption, and WS-Security in 2026 covers the message layer on a modern JDK.