Kubernetes Networking: How a Service Reaches a Pod
A Kubernetes Service has an IP address that no machine owns and no process listens on. Yet when you send traffic to it, the traffic arrives at one of your Pods. This post explains how, one step at a time.
Click the steps in the diagram below, in order. On the left are kubectl, outside the cluster, and the control plane: the API server, etcd and the EndpointSlice controller, where the cluster’s records are made and kept. On the right is a worker node, where the Pods run and the traffic actually flows. Each step plays a numbered list of actions. Each action names who does it, the exact API call, watch event, etcd write or packet involved, and what changes as a result. The grey box on each card shows what that component holds or is doing at that moment. Watch the etcd card fill up with the actual keys.
The problem a Service solves
Pods are disposable. When a Pod crashes, is rescheduled, or is replaced during a rollout, the new Pod gets a new IP address. If a frontend talked to a backend by Pod IP, it would break every time the backend changed.
A Service fixes this by giving a group of Pods one stable address. Clients talk to the Service. Kubernetes keeps track of which Pods are behind it right now, and quietly sends each connection to one of them.
The interesting part is how it does that, because the answer is not what most people expect. There is no proxy process in the middle of the traffic. There is no load balancer box. The forwarding is done by the Linux kernel on each node, using rules that were written ahead of time.
Step 1: the Pods start
The scheduler has placed two Pods on the node. The kubelet asks the container runtime to start their containers. Then the CNI plugin (Container Network Interface) gives each Pod its own IP address on the cluster network:
- Pod A:
192.168.1.10 - Pod B:
192.168.1.11
Both Pods carry the label app=backend. Labels are plain key-value tags. They mean nothing on their own, but they are how a Service finds its Pods.
These IPs are real and routable inside the cluster. Any Pod can reach 192.168.1.10 directly. They are also temporary. Delete Pod A and its replacement will get a different address.
Step 2: the Service is created
Someone applies a Service manifest:
apiVersion: v1
kind: Service
metadata:
name: backend-svc
spec:
selector:
app: backend
ports:
- port: 80
targetPort: 8080
The API server validates it, then allocates a ClusterIP from the cluster’s Service IP range (for example 10.96.0.0/12). Here it picks 10.96.0.50. The Service, with its IP, is saved in etcd.
Notice what has not happened. No network interface has this address. No process is listening on port 80 at 10.96.0.50. If you could ping it, nothing would answer. The ClusterIP is a virtual IP. At this point it is just a value in a database record.
The Service also says which Pods it stands for, through its selector: every Pod labelled app=backend. But it does not list them. Finding them is someone else’s job.
How the API server picks the ClusterIP
No controller and no outside service is involved. The kube-apiserver picks the address itself, while it handles the create request, before the Service is saved. That is why the 201 Created response already contains the ClusterIP.
1. The range comes from a flag. kube-apiserver is started with --service-cluster-ip-range, for example 10.96.0.0/12, kubeadm’s default. Every ClusterIP comes from this block. It must not overlap with the Pod network or the node network. A few addresses are taken by convention:
- The first address,
10.96.0.1, belongs to the built-inkubernetesService in thedefaultnamespace. Pods use it to reach the API server. 10.96.0.10is usually the cluster DNS Service. kubeadm sets that up; it is a convention, not a rule.
2. It chooses an address.
- If the manifest sets
spec.clusterIP, the API server checks that the address is inside the range and not already taken, and uses it. - If it sets
clusterIP: None, the Service is headless: no IP is allocated, and DNS returns the Pod IPs directly. - Otherwise it picks a random free address. The range is split in two: a small lower part kept for people who choose an IP by hand, and the rest for random picks. Random picks use the upper part first, so automatic Services rarely take an address someone wanted to set by hand.
3. It reserves the address. Production control planes run several API server replicas. Two of them must never hand out the same IP, so the reservation is written to etcd in a way that only one can win.
- Recent Kubernetes versions describe the range as a ServiceCIDR object. To reserve
10.96.0.50, the API server creates an IPAddress object whose name is10.96.0.50, with a reference back to the Service. Object names are unique, so if another replica took that IP a moment earlier, the create fails and the API server tries another address. You can see both withkubectl get servicecidrsandkubectl get ipaddresses. When a range fills up, you can add another ServiceCIDR while the cluster runs. - Older versions kept the whole range as one bitmap of used and free addresses, stored in a single etcd key,
/registry/ranges/serviceips. Reserving an IP meant flipping one bit and writing the key back with a compare-and-swap. If another replica had changed it first, the write was rejected and retried.
4. It saves the Service. The address goes into spec.clusterIP and the Service is written to etcd. If that write fails, the reservation is released.
5. It cleans up afterwards. A repair loop inside kube-apiserver regularly compares the reserved addresses with the existing Services. It frees reservations left behind, for example by a crash between steps 3 and 4. It also reports Services whose IP has no reservation.
So creating one Service writes two things to etcd: the IP reservation and the Service itself.
A dual-stack Service goes through this once per IP family and gets one IPv4 and one IPv6 address. A NodePort or LoadBalancer Service also gets a node port the same way, from a separate port range (--service-node-port-range, by default 30000 to 32767).
Step 3: the controller finds the Pods
Inside the controller-manager runs the EndpointSlice controller. Like every controller in Kubernetes, it watches the API server and closes gaps between what is wanted and what exists.
It sees a new Service with the selector app=backend. It looks up every Pod with that label, keeps only the ones that are ready, and writes their IPs into a separate object called an EndpointSlice:
apiVersion: discovery.k8s.io/v1
kind: EndpointSlice
metadata:
name: backend-svc-x7k2p
labels:
kubernetes.io/service-name: backend-svc
addressType: IPv4
ports:
- port: 8080
endpoints:
- addresses: ["192.168.1.10"]
conditions: { ready: true }
- addresses: ["192.168.1.11"]
conditions: { ready: true }
You can see it yourself with kubectl get endpointslices -l kubernetes.io/service-name=backend-svc.
This controller keeps working forever. When a Pod fails its readiness probe, it is removed from the slice. When a new Pod with the label becomes ready, it is added. The Service object never changes. Only the EndpointSlice does.
The “ready” filter is important. A Pod that is still starting, or failing its readiness probe, does not receive traffic. That is how rolling updates avoid sending requests to Pods that are not ready to serve them.
Why a separate object? A large Service can have thousands of Pods behind it. Splitting the list into slices of up to 100 endpoints by default means a single Pod change only rewrites one small slice, instead of one giant list that every node has to download again.
Step 4: kube-proxy programs the kernel
So far, everything is just records in etcd. Nothing on the node knows how to reach 10.96.0.50. That is the job of kube-proxy.
kube-proxy runs on every node in the cluster. It opens a long-lived watch on the API server, an HTTPS request that stays open and streams changes as they happen, for Services and EndpointSlices. It never talks to etcd directly. The moment the EndpointSlice from step 3 appears, kube-proxy receives it.
kube-proxy then translates those records into NAT rules in the Linux kernel. In the default iptables mode, the rules for our Service look roughly like this (names shortened):
# Anything going to the ClusterIP on port 80 jumps to the Service chain
-A KUBE-SERVICES -d 10.96.0.50/32 -p tcp --dport 80 -j KUBE-SVC-BACKEND
# Pick an endpoint: first one with probability 0.5, otherwise the second
-A KUBE-SVC-BACKEND -m statistic --mode random --probability 0.5 -j KUBE-SEP-PODA
-A KUBE-SVC-BACKEND -j KUBE-SEP-PODB
# Rewrite the destination to the chosen Pod
-A KUBE-SEP-PODA -p tcp -j DNAT --to-destination 192.168.1.10:8080
-A KUBE-SEP-PODB -p tcp -j DNAT --to-destination 192.168.1.11:8080
On a node you can see the real ones with sudo iptables-save -t nat | grep backend-svc.
Read it top to bottom. A packet for 10.96.0.50:80 is sent to the Service’s chain. That chain rolls a die: half the time it goes to Pod A, otherwise to Pod B. With three Pods, the probabilities would be 1/3, then 1/2, then the rest, which works out to an equal share for each. The last rule does the actual work, DNAT (destination NAT): it rewrites the packet’s destination address from the ClusterIP to the Pod’s IP, and the port from 80 to 8080.
Here is the key insight. kube-proxy is not in the traffic path. It is a configuration agent. It writes the rules and then gets out of the way. If you stopped kube-proxy right now, existing rules would keep forwarding traffic. It would only stop updating them when Pods come and go.
Because every node runs kube-proxy, every node ends up with the same rules. A client on any node can reach the Service, even if none of the backend Pods run on that node.
iptables, nftables and IPVS
kube-proxy has several modes that do the same job with different kernel features:
- iptables is the long-standing default. Simple and reliable, but rules are checked in order, so with many thousands of Services lookups and rule updates get slower.
- nftables is the modern replacement for iptables in the Linux kernel. It uses lookup tables (maps) instead of long rule lists, so it scales much better, and it is the direction upstream Kubernetes is moving.
- IPVS is a load balancer built into the kernel, with a choice of algorithms such as round-robin or least-connections. It was the traditional answer for large clusters, though newer clusters are encouraged to use nftables instead.
Some network plugins, such as Cilium, can replace kube-proxy entirely and do the same translation with eBPF programs in the kernel. The idea stays the same: the decision is made in the kernel, on the node, from rules computed in advance.
Step 5: traffic flows
Now a client Pod on the node sends a request to 10.96.0.50:80. Watch what happens in the widget: the control plane goes dark, and only the kernel and the Pod light up.
- The packet leaves the client Pod with destination
10.96.0.50:80. - Before it is routed, the kernel checks its NAT table. The
KUBE-SERVICESrule matches. - The random rule picks Pod A.
- DNAT rewrites the destination to
192.168.1.10:8080. - The kernel routes the packet to Pod A like any other packet. If Pod A were on another node, the CNI network would carry it there.
The API server, etcd, the controller-manager and even the kube-proxy process do nothing during this. They already did their part. The data path is just the kernel following rules. This is why a Service adds almost no latency, and why traffic keeps flowing even if the control plane goes down for a while.
What about the replies?
Pod A sees a packet from the client and answers it. But the client sent its request to 10.96.0.50, not 192.168.1.10. If the reply came back from the Pod’s own IP, the client would not recognise it.
The kernel handles this with connection tracking (conntrack). When it rewrote the first packet, it recorded the translation. Every reply on that connection is automatically rewritten back, so it appears to come from 10.96.0.50:80. Every later packet on the same connection goes to the same Pod, without rolling the die again. The load balancing happens per connection, not per packet or per request.
That last point has a practical consequence. A client that opens one long-lived connection, such as HTTP/2 or gRPC, will send all its requests to a single Pod. Spreading those requests evenly needs load balancing at the application level, not just a Service.
How clients find the ClusterIP
Nobody writes 10.96.0.50 into their code. The cluster runs a DNS server, usually CoreDNS, which watches Services just like kube-proxy does. Every Service gets a name:
backend-svc.default.svc.cluster.local -> 10.96.0.50
Pods in the same namespace can simply use backend-svc. So the full path of a request is: DNS turns the name into the ClusterIP, and the kernel turns the ClusterIP into a Pod IP.
When a Pod goes away
The whole chain repeats on its own whenever something changes. Say Pod A is deleted:
- Pod A is marked as terminating. The EndpointSlice controller marks it as not ready in the slice.
- kube-proxy on every node receives the updated slice through its watch.
- It rewrites the kernel rules so that only Pod B is left.
- New connections all go to Pod B. When a replacement Pod becomes ready, it is added back the same way.
No client had to be told anything. The Service name and the ClusterIP never changed.
There is a short window between steps 1 and 3 while the update spreads to every node. This is why well-behaved applications keep serving for a few seconds after they receive the signal to shut down, often with a small preStop delay, so they do not drop requests that are already on their way.
The pattern behind it
This flow is the same pattern as the rest of Kubernetes, described in the cluster architecture post:
- Everything goes through the API server. The Service, the EndpointSlice and every update to them are stored in etcd through it.
- Nobody pushes, everybody watches. The controller watches Services and Pods. kube-proxy watches Services and EndpointSlices. CoreDNS watches Services. None of them call each other.
- The control plane decides; the nodes do. The control plane only ever writes records. The work of moving packets happens entirely on the nodes, in the kernel.
Quick recap
| Piece | Where it runs | Its job in this flow |
|---|---|---|
| CNI plugin | Every node | Gives each Pod a real, routable, temporary IP |
| API server + etcd | Control plane | Allocates the ClusterIP and stores the Service and EndpointSlices |
| EndpointSlice controller | Control plane (controller-manager) | Matches the selector to ready Pods and keeps their IPs listed |
| kube-proxy | Every node | Watches Services and EndpointSlices, writes NAT rules into the kernel |
| Linux kernel (iptables / nftables / IPVS) | Every node | Picks a Pod per connection and rewrites the packet’s destination |
| conntrack | Every node, in the kernel | Keeps a connection on one Pod and rewrites replies back |
| CoreDNS | Cluster add-on | Turns backend-svc into the ClusterIP |
This post was written by an AI agent. It represents my understanding of how traffic reaches a Pod through a Kubernetes Service.
Mahmoud Elshenhab