Delayed Feedback in Distributed Routing
Stale load reports can send multiple routers toward the same destination. DLB illustrates how probability updates and control intervals interact, and what its production comparison establishes.
When observations arrive late, routing requires a decision about how quickly to change traffic shares as well as where to send requests.
Delayed load observations
Sending every request to the least busy destination does not guarantee balanced load across multiple routers. If the routers act on the same past state, their new requests accumulate at the same destination. The next report may make another destination look idle, moving the concentration of requests again.
Google's Distributed Load Balancing (DLB) system addresses this problem in generative model inference. The DLB paper separates roots, which choose destination cells, from leaves, which choose inference servers within a cell. This article uses the arXiv v1 dated September 17, 2026, available as of September 2026.
A cell is a routing unit that groups inference servers. Its leaf reports Requests in Flight (RIF), covering queued and executing requests, together with latency models and error rates. Roots use these reports and network latency to evaluate destination costs.
A leaf's report already describes a past state when a root makes its decision. Requests from other roots arrive in the meantime, so the lowest observed cost may differ from the lowest cost at arrival. The diagram reconstructs the request and observation paths in Section 3.1 of the paper.
Routing by request characteristics
DLB describes discrete routing, which sends each request directly to the cell with the lowest observed cost. It uses this approach for sparse, heavy requests called boulders, where each placement matters. For short requests arriving at high rates, called sand, the rate at which placement decisions accumulate also matters.
Flow routing maintains a probability for each destination and updates those probabilities gradually. Requests follow the current distribution, so a changed observation does not immediately move all traffic to one cell. Section 3.2 connects this distinction to oscillations caused by delayed observations.
| Dimension | Discrete routing | Flow routing |
|---|---|---|
| Decision variable | Lowest-cost cell for this request | Routing probability for each destination |
| Response to a new observation | Immediately affects cell selection | Gradually adjusts probabilities |
| Request class in the paper | Sparse, heavy boulders | Short, frequent sand requests |
| Main consideration | Cost of an individual placement | Simultaneous reactions and oscillation across roots |
Request count alone does not determine this distinction. The work represented by each request and the load observation interval together determine how many placement decisions accumulate before the next observation.
Probability updates and control intervals
Flow routing subtracts a cost-proportional adjustment from the current probabilities and projects the result onto the permitted probability set. Projection keeps probabilities nonnegative and makes their sum equal to one. It also keeps probabilities at zero for cells excluded by restrictions such as data residency policies.
The notation below transcribes Equation (2); it is not executable code. Destination costs use load observations that arrive after a delay.
x_i(t + delta_t) = Pi_Delta_i(x_i(t) - eta_i * delta_t * c_i(t))| Symbol | Meaning | Condition to inspect |
|---|---|---|
| x_i | Destination probability vector for root i | Assign probability only to permitted cells |
| c_i | Vector of observed destination costs | Difference between current and observed load |
| eta_i | Update step size | Strength of the response to cost changes |
| delta_t | Probe update interval | Time between observations and updates |
| Pi_Delta_i | Projection onto the permitted probability set | Nonnegative values, sum of one, allowed routes |
The update includes the product of the step size and the interval. Comparing step sizes without recording the interval leaves the size of each probability adjustment unclear. The interval also changes observation frequency, so both values need to be recorded together.
Adaptation speed and stability
A larger step size reacts more quickly to a change in cost. It also makes larger probability changes based on older observations, which can increase errors caused by delay. Section 4 analyzes this relationship with a fluid model that approximates request arrivals and processing as continuous flows.
| Control or analysis target | Expected effect or measurement | Condition to retain |
|---|---|---|
| Larger update step | Faster adaptation to load changes | Excessive reaction to delayed costs |
| Smaller update step | Smaller probability movement per update | Recovery speed after a change |
| Theoretical stability | Bound on time-averaged deviation | Fixed delays and assumptions on service rates, capacity, and costs |
| Production latency | Response delays experienced by requests | Traffic composition and routing layer |
The analysis concerns time-averaged deviation, which is a different quantity from an upper percentile of response latency. Results under fixed-delay and service-rate assumptions do not establish a latency guarantee for the entire production system. In particular, they should not be extended to every discrete routing policy.
Interpreting the production comparison
Section 5.2 analyzes 68 endpoints that migrated from the previous router to DLB. It keeps Prequal, which selects servers within each cell, unchanged. The analysis adjusts for demand and endpoint effects to evaluate the change in routing across cells. The following reductions are estimates from that regression analysis.
| Metric | Estimated reduction relative to the previous router | Scope |
|---|---|---|
| p50 latency | About 17% | Median latency of the analyzed endpoints |
| Mean latency | About 13% | Mean under the same comparison conditions |
| p90 latency | About 14% | 90th percentile of the latency distribution |
| p95 latency | About 13% | 95th percentile of the latency distribution |
This analysis observes an actual migration; it is not a randomized controlled experiment. Its percentages therefore cannot be presented as expected gains for other models or clusters. This article does not execute the paper's code or reproduce its traffic and results.
Observations before adoption
Applying the approach to another service starts with recording the observation path separately from the request path. Recording when load was measured and when a router used it separates routing decisions from information freshness. The following operational checks are derived from the structure described in the paper.
| Observation | Question to answer |
|---|---|
| Load measurement and receipt timestamps | How old is the observation used by the router? |
| Destination probability changes | Do several roots move together after the same observation? |
| Queued and executing requests, processing latency | Do similar request counts represent different amounts of work? |
| Permitted destination set | Do probability updates preserve placement constraints? |
| Recovery after a load change | Do latency and request distribution stabilize together? |
These records make it possible to interpret changes after adjusting the step size. Smoother destination probabilities and lower user-facing latency also need separate verification. The existing load balancing article covers the basic algorithm families.
Summary
Distributed routing includes the problem of multiple decision makers acting on delayed observations. DLB's flow routing gradually changes destination probabilities and treats the step size and update interval together. Theoretical stability and production latency improvements require validation under their respective assumptions and metrics. Adoption should connect observation age, probability changes, and actual response latency.