Kubernetes readiness probes can mislead during updates; understanding true application readiness is key for seamless deployments.

Understanding Kubernetes Readiness Probes
Kubernetes rolling updates are celebrated for their potential to deploy applications without downtime. However, the reality can be quite different if readiness probes are not configured properly. These probes are intended to signal when a pod is ready to handle production traffic, but a prevalent issue is that they may only confirm process health rather than actual application readiness.
The Discrepancy in Results
Take, for example, a situation with a 5G signaling service. During a recent rolling update, readiness probes returned positive results; new pods appeared healthy and the old ones drained gracefully. Yet, amidst this seemingly flawless transition, an 18% spike in session-establishment failures was noted.
The underlying mechanism here is critical: the new pods might pass standard HTTP checks, but they could still be unprepared to manage traffic. This scenario reveals a significant flaw in default readiness probe configurations commonly used in Kubernetes deployments.
What Readiness Probes Are Missing
Standard Kubernetes readiness probes primarily utilize HTTP checks, which simply involve sending a GET request to a specified endpoint. A 200 response leads Kubernetes to declare the pod ready and commence traffic routing. While effective for stateless applications, this approach falls short when workloads require more than a simple HTTP response.
In the case of telecom functions, protocols like SIP, Diameter, or GTP necessitate that new pods complete a series of registrations and handshakes with upstream systems. This process may take longer than the probe response time, resulting in a misleading state where the pod is signaled as ready while still in the process of establishing essential connections.
Why the Problem Persists
For telecom functions, and indeed for any stateful service, the registration and peer relationship establishment can delay readiness significantly, sometimes exceeding a minute post-startup. During this span, even though Kubernetes sees a healthy pod and begins routing traffic, the application hasn't completed critical steps necessary to serve requests effectively.
Consequently, instead of outright failures, users may experience subtle degradations. Increased latency, failed session attempts, and a perplexing absence of errors on monitoring dashboards can obscure the real issues at play. This misalignment between Kubernetes’ oversight and the actual operational requirements leads to what many in the industry are starting to recognize as a hidden failure during rolling updates.
Case Study: The 5G Update Failure
In the previously mentioned 5G application rollout, the default HTTP probe was configured to check against a basic endpoint. It confirmed that the pod had started and returned a 200 status within mere seconds. However, the true readiness of the pod—its ability to handle traffic post-SIP registration—was compromised by this simplistic health check.
The resulting session establishment issues became evident, particularly during high traffic periods. The quick probe succeeded in passing checks, but the full implications of the pod’s unpreparedness were only apparent after the fact, allowing for the impact of the service disruption to go unnoticed initially.
Implementing Effective Protocol-Aware Readiness Probes
Shifting from a traditional HTTP readiness probe to a protocol-aware approach can significantly mitigate these challenges. For the signaling service, a redesign of the probe to include a script that actively verifies registration status can make all the difference. Such a probe would look like this:
readinessProbe:
exec:
command:
- /bin/sh
- -c
- “check-registration-status.sh && exit 0 || exit 1”
initialDelaySeconds: 10
periodSeconds: 5
failureThreshold: 12
This configuration allows the probe to wait until the pod confirms it has established its necessary connections before it gets marked as ready. The script becomes essential—holding back traffic until the pod is genuinely prepared to serve requests.
Conclusion: Addressing the Gap Between Running and Ready
The distinction between a pod being “running” versus “ready” is a crucial one for Kubernetes users. Out-of-the-box configurations may work well for automated checks, but they can lead to blind spots, especially in stateful or protocol-dependent environments.
Ultimately, the readiness probe must encapsulate what it truly means for an application to serve real traffic—going beyond a simple check for running processes or basic HTTP health. By ensuring that your readiness probes verify essential application conditions, teams can maintain both operational integrity and higher reliability during critical updates.
Discussion
Sign in to join the discussion.