On Wed, Sep 30, 2026 at 11:33 PM Willy Tarreau <[email protected]> wrote:
> On Wed, Sep 30, 2026 at 10:27:27PM -0700, David Birdsong wrote: > > Concretely: take a multi-tenant backend with, say, 200 application > servers, > > one hash key per tenant. Today, with map-based or consistent hashing, a > > tenant's traffic is deterministic under steady state, but there's no > fixed > > membership guarantee once servers churn. With consistent hashing > > specifically, when a server goes down the ring walk moves that tenant's > > traffic to whatever server is next around the ring -- which is a > function of > > everyone else's current health, not a fixed property of the tenant's key. > > Over enough failures or scaling events, a tenant's traffic can, in the > > worst case, land on any server in the backend. There's no way to answer, > > ahead of time, "which servers can tenant X's requests ever reach" with > > anything narrower than "the whole backend, eventually." > > You've mentioned this point before but it's not clear to me: what do > you call a "multi-tenant backend" ? You seem to imply that you would > place multiple applications inside the same backend, but that's strange > because a backend currently *is* one instance of an application, with > its own rules, load balancing, stickiness etc. That's why I'm not sure > that's what you mean and am still confused about your usage of "tenant" > here. > > > That's the problem I want to solve: containment. If a tenant is noisy -- > a > > bug, an abusive client, a runaway batch job -- its blast radius today is > > effectively unbounded. It can, over time, degrade service for every other > > tenant sharing the backend, because nothing enforces a ceiling on how > many > > distinct servers it can touch. > > It entirely depends on the LB algorithm. If you're using round-robin, > of course, all servers will be visited. With leastconn, possibly less > but still many. With hashing, it will solely depend on the hashed > criteria and how the client can control them (and their willingness to > intentionally cause harm of course). For example, consistent-hashing is > used a lot with caches because it maintains a high cache hit ratio by > sending the same URL to the same server, which means that a client > hammering a given URL will not affect other nodes (but the client that > scans many URLs will). And combined with the load-factor, it will make > use of adjacent nodes for the same URL, spreading the load on the > smallest subset needed to handle the load. Other services will hash > on the client address to maintain a form of rough stickiness and avoid > cache reloads most of the time. > > > With rendezvous-subset and, say, hash-candidates 3, each tenant key is > > ranked against all 200 servers up front, independent of health, and is > > permanently bound to its top 3. > > This is precisely the point I just cannot grasp based on my > misunderstanding of what you call a tenant in this context. > > > Health and load only ever pick among those > > 3 (or let us queue/redispatch within them); a tenant literally cannot > reach > > a 4th server, no matter what else happens in the backend. So "which > servers > > can tenant X ever reach" becomes a static, computable fact -- three named > > servers -- instead of "potentially all of them, given enough churn." > That's > > the property the ring walk's health-dependent fallback can't give at any > > virtual-node count, which is why I think it needs a new algorithm rather > > than a tweak to consistent hashing. > > At this point I got totally lost again :-/ > Say we have 100 servers, srv1..srv100, and a vhost acme.example.com that maps to one customer which I've been calling a tenant. What you're picturing (plain hashing, affinity/caching): hash something about the request — URL param, IP, header — to pick one server, mainly for cache locality. If that server's down, we fail over into the rest of the backend; any of the other 99 can end up serving it. The hash picks a starting point, not a boundary — there's no limit on how much of the backend a given key can eventually reach. BTW, I requested consistent-hashing back in 2009 and offered some sponsorship. I've been using it heavily since then. Many thanks! What we want instead: give each vhost a fixed, small slice of the 100 servers — not for cache locality, but as a resource/isolation boundary. A vhost should never be able to use more than, say, 3 servers' worth of capacity, no matter what else is happening in the backend, and we should be able to name exactly which 3 ahead of time. A side effect is some affinity cache locality properties that happen to work well for resource bin-packing and helps to offset the downsides I explain farther down. How rendezvous-subset does it: rank all 100 servers against the vhost's key. That gives a fixed, deterministic order, e.g. for acme.example.com: srv37, srv81, srv52, srv6, srv93, ... (all 100, ranked) With hash-candidates 3, that vhost is confined to the top 3: srv37, srv81, srv52. Those are the only servers that vhost will ever be assigned to — forever, since the ranking only depends on the vhost's key and server identity. The only thing that changes this list is N changing (a server added to or removed from the backend). Health, load, maintenance — none of that reorders or extends it. If we raised Y to 4 for this vhost, they'd get the same first 3 plus one more (srv6) — Y just slides further down the same fixed ranking, it doesn't recompute a different set. Picking which one of those 3 servers handles a given request is a separate, second decision. For simplicity, let's say leastconns for now — but that choice never changes the subset itself. In my prototype, I have some very interesting, configurable choosing mechanisms within that subset I can share once the subset selection how and why is clear. The sharp edge, and the whole point: if all 3 of a vhost's servers are down, in maintenance, or saturated (at maxconn with no queue room left), that vhost gets a 503. We do not fail over into the other 97 or enqueue on the proxy queue. The backend as a whole can be perfectly healthy with capacity to spare, and one vhost can still be fully down — because its blast radius is capped at 3 servers by design, not by luck. For further background: this was inspired by AWS's shuffle sharding. > > Willy >

