TCP-over-TCP Is a Bad Idea
Encapsulating a TCP stream inside another TCP connection destroys the feedback mechanisms both rely on to manage congestion and reliability. The inner and outer sessions each maintain independent sequence spaces, ACK processing, and congestion-control state. Their control loops interfere destructively.
TCP interprets any delay in ACK reception as potential congestion. Its cwnd (congestion window) growth follows additive-increase/multiplicative-decrease (AIMD) based on ACK pacing and loss events (duplicate ACKs or retransmission timeouts). When the outer TCP link drops a segment or reorders packets, the inner TCP never sees actual packet loss, it sees delayed delivery. That delay withholds ACKs, triggering spurious retransmissions and cwnd collapse inside the tunnel.
Because the outer layer enforces in-order delivery, one lost outer segment stalls the entire byte stream. Until a missing segment is retransmitted and acknowledged (driven by the outer layer’s RTO and fast-retransmit logic), all subsequent bytes are held in the sender’s buffer. The inner TCP session thus experiences an artificial “black hole” where the network appears to stop forwarding packets.
Typical sequence:
- Outer TCP sends segments with SEQ=1000–2000, 2000–3000, etc. Segment 2000–3000 is lost.
- Receiver sends duplicate ACKs for SEQ=1000. Outer TCP triggers fast retransmit after three dup-ACKs, sends retransmit with SEQ=2000.
- During this interval, no new payload is delivered upward. The inner TCP’s send queue fills, ACKs from the receiver stall, RTO expires.
- Inner TCP retransmits and halves its cwnd. Outer TCP, already retransmitting, also halves its cwnd. Throughput collapses quadratically.
The interaction between delayed ACKs (from the outer) and retransmission backoff (from the inner) causes exponential latency growth. Both layers may increase their RTO independently, compounding recovery time. Window scaling and flow-control interactions make it worse. If the outer layer’s advertised window closes due to full buffers, the inner sender perceives it as path congestion.
TCP flags add no rescue here.
- ACK and PSH flags from the inner session are meaningless until the outer TCP drains its own send queue.
- RST or FIN on the inner stream can be delayed or even reordered by the outer layer’s retransmission behavior, corrupting connection-teardown semantics.
- URG data is treated as normal payload inside the outer stream and loses urgency entirely.
Multiplexing several inner connections over one outer TCP magnifies head-of-line blocking. A single missing segment halts delivery to every stream, regardless of which inner session originated it.
Instrumentation lies. The inner connection’s RTT samples (from TCP timestamps or SACK timing) now measure tunnel buffering delay, not network delay. The outer connection’s loss recovery obscures actual path characteristics. Congestion algorithms such as CUBIC or BBR on the inner layer receive useless signals and miscompute bandwidth estimates.
This configuration defeats the purpose of TCP’s congestion avoidance: coordinated, end-to-end rate control tied to actual packet loss and delay. The two independent feedback loops create positive feedback instability.
If tunneling is required, the correct approach is a UDP-based carrier (e.g., QUIC, WireGuard, or custom encapsulation). UDP preserves loss visibility and packet boundaries so the inner transport’s control loop sees genuine network behavior.
If TCP-over-TCP is unavoidable, constrain it to a single stream, disable Nagle’s algorithm (TCP_NODELAY), cap buffers to reduce hoarding, and accept degraded performance. But understand: this is a stopgap, not transport design.
TCP assumes it is the transport. Wrapping it in another TCP corrupts every signal it uses to function: sequence, timing, loss, and flow control. The protocol can’t correct for that.
Member discussion