4 min read

How Broken PMTU Causes Certificate-Based Failures in Cloud Paths

How Broken PMTU Causes Certificate-Based Failures in Cloud Paths
Photo by Ricardo Gomez Angel / Unsplash

When certificate-based traffic starts failing intermittently after a move to the cloud, latency usually gets blamed first.

That is understandable. The path is now longer, round-trip times are higher, and the failures often show up in flows that already involve multiple exchanges: TLS handshakes, mutual TLS, EAP-TLS, RADIUS-backed authentication, API calls with client certificates, reverse proxies doing certificate validation, and other certificate-heavy protocols.

The symptoms fit the story well enough. Handshakes stall halfway through. Some clients succeed while others fail. Smaller certificate chains work more often than larger ones. Peak periods make the whole thing look worse.

Sometimes it really is latency.

But a lot of the time it is broken PMTU discovery somewhere along the path.

The Pattern

The setup is usually not reckless. Certificates are valid. Trust stores are correct. Cipher and protocol settings line up. The cloud migration itself is not obviously wrong. Connectivity tests pass. Basic functional testing passes too.

Then production traffic starts exposing a different class of problem.

What people see tends to look like this:

  • TLS or certificate-based handshakes timing out mid-exchange
  • larger certificate chains failing more often than smaller ones
  • intermittent failures across clients, services, or sites
  • retransmissions and retries making the issue look like instability under load

That last part matters. These problems rarely present as a clean hard failure. They usually show up as partial success, sporadic breakage, and a lot of wasted time arguing about whether the problem lives in the application, the PKI, the load balancer, the firewall, or the cloud provider.

Why latency is not a satisfying explanation

One clue comes up again and again: the problem improves when packet size is reduced.

That can happen in different ways. Someone reduces an application-layer fragment size. Someone trims a certificate chain. Someone changes a handshake-related sizing parameter. Someone removes a certificate that did not strictly need to be sent. Suddenly the failures become less frequent or disappear entirely.

At that point, “latency” stops being a complete explanation.

If the same amount of data is split into smaller chunks, more packets are needed to carry it. If round-trip delay were the core problem, needing more packets would not obviously help. It might still interact with timer behavior, but it does not explain why the flow becomes stable once the packet sizes drop.

A packet-size problem explains that much better.

What is actually happening

In a healthy path, when a packet is too large for a downstream hop, the sender gets feedback and adjusts. That is the normal PMTU process. The sender learns that the path cannot carry packets above a certain size and retransmits at a lower size.

When that process breaks, the path starts behaving badly in a very specific way.

A packet is sent at a size that some hop cannot carry. The packet gets dropped. The sender does not receive the feedback it needs to reduce size. It sends the same thing again. From the point of view of the protocol above it, the exchange just looks slow, unstable, or random.

That is how PMTU failures get misread as latency problems.

The path is not merely slower than expected. It is silently failing for traffic above some size threshold.

Why certificates expose this so often

Certificates make the problem easier to trigger because they increase handshake size.

A shallow chain with compact certificates may stay comfortably under the path’s real limit. A deeper chain, larger certificate, stapled response, or extra handshake payload may push the exchange over it. That is why these incidents often look selective at first.

Some clients work. Some clients with larger chains do not. Some paths succeed. Others fail. Nothing looks deterministic unless you look at packet sizes.

This is also why people end up blaming the certificate chain itself. The chain is not necessarily wrong. It is simply large enough to expose a path-size problem that smaller exchanges do not hit.

Why timeouts sometimes seem to fix it

Timeout increases can improve success rates, but that does not mean latency was the original cause.

Once a flow starts suffering drops, retries, and retransmissions, it naturally takes longer. Extending a timeout gives the system more room to survive the failure pattern. That can be operationally useful, but it is not the same as fixing the path.

The same applies to reducing handshake or fragment sizes. Those changes can be perfectly reasonable mitigations. But they are often workarounds around a broken PMTU path, not proof that the network is otherwise healthy.

Where this tends to hide

The problem is rarely visible in a simple reachability check. A path can be routable, encrypted, and mostly functional while still mishandling larger packets.

Once traffic crosses a more complex cloud path, there are more places where this can happen:

  • load balancers
  • firewalls
  • virtual appliances
  • overlay networks
  • tunnels
  • VPN edges
  • proxies
  • peering boundaries
  • any filtering policy that treats ICMP as noise

That is enough to make PMTU failures surprisingly persistent, especially when different teams own different segments of the path.

What to check

If certificate-heavy traffic starts failing only in some cases, packet size belongs high on the list immediately.

The useful questions are straightforward:

  • Do larger certificate chains fail more often than smaller ones?
  • Does reducing handshake or fragment size improve reliability?
  • Do failures correlate with specific paths, proxies, or cloud edges?
  • Are there signs of repeated retransmission at the same size?
  • Is the path handling PMTU feedback correctly end to end?

The fastest way to answer those questions is still packet capture. Ideally, capture at both ends and at any intermediary point you control. If the larger packets are not arriving, or if retransmission keeps happening at the same size, the picture gets clearer very quickly.

What usually helps

The practical mitigations are familiar:

  • reduce handshake or fragment size where the protocol allows it
  • simplify certificate chains where appropriate
  • remove unnecessary intermediates or excess handshake payload
  • review devices and policy controls along the path
  • verify that PMTU-related feedback is not being filtered or ignored
  • tune timeouts only with a clear understanding of whether they are compensating for packet loss

Those steps are all reasonable. What matters is understanding whether you are fixing the root cause or stepping around it.

The important part

If a certificate-related failure improves when you make packets smaller, do not stop at “the cloud added latency.”

That may be the symptom people noticed first. It is often not the real mechanism.

A surprising number of intermittent TLS and certificate failures in cloud environments come down to broken PMTU discovery somewhere along the path, usually made worse by the fact that certificate-heavy exchanges are exactly the kind of traffic that expose hidden size limits first.

The network does not have to be fully broken for this to happen.

It only has to be broken above a certain packet size.