Architecture¶
How the LoxiLB Inference Gateway sits in an LLM serving stack, and how its control plane and data plane divide the work of inference-aware routing.
The serving path¶
Clients speak OpenAI-compatible HTTP (and SSE for streaming) to a VIP on the gateway.
The L7 fullproxy (mode: 4) terminates the connection, inspects the request, selects a
model pool and endpoint, and proxies to vLLM, SGLang, TensorRT-LLM, or llama.cpp. The
engine determines the optional cache-event and P/D contract; the four engines are not
interchangeable.
flowchart LR
CLIENT([OpenAI-compatible<br/>HTTP or SSE client]) --> FP
subgraph GW [LoxiLB Inference Gateway]
API["Management API<br/>/netlox/v1"] --> RULES["Validated rules<br/>and endpoint state"]
RULES --> EBPF["eBPF L4 data path"]
RULES --> FP["Fullproxy L7 data path<br/>HTTP parsing and routing"]
EVENTS["KV inventory services<br/>tokenizer and engine adapters"] --> FP
end
FP --> V["vLLM<br/>ZMQ events; sequential P/D"]
FP --> S["SGLang<br/>ZMQ per DP rank; concurrent P/D"]
FP --> T["TensorRT-LLM<br/>HTTP event drain; sequential P/D"]
FP --> L["llama.cpp<br/>plain pool; no KV-exact or P/D"]
V -. block hashes .-> EVENTS
S -. block hashes .-> EVENTS
T -. destructive event drain .-> EVENTS
style GW fill:#e1f5fe,stroke:#0288d1
style EVENTS fill:#e8f5e9,stroke:#43a047
style L fill:#fff3e0,stroke:#f57c00
The gateway is a single Go/eBPF binary. There is no Envoy, no ext-proc sidecar chain, and no mandatory Kubernetes control plane — L4 through inference-aware L7 lives in one process.
Control plane vs. data plane¶
The gateway splits cleanly into a control plane and a data plane.
Control plane — a GoLang process that owns configuration and routing intelligence:
- The REST API listens on port 11111 at
/netlox/v1/.... Load balancers are created and listed at/config/loadbalancerand/config/loadbalancer/all. This is where every rule, endpoint, API key, and policy is programmed. - The Go KV inventory services consume vLLM/SGLang ZMQ events or TensorRT-LLM's HTTP event drain and maintain per-rule, per-endpoint block-hash inventories. llama.cpp has no supported KV-event plane.
- The control plane validates engine/topology combinations and programs the rule into the packet and fullproxy data paths.
Data plane — where packets and bytes actually move:
- eBPF handles the L4 fast path (NAT modes, connection tracking) for classic load balancing, inherited unchanged from upstream loxilb.
- The sockproxy / fullproxy is a userspace HTTP proxy. When a rule runs in
mode: 4, it terminates the client TCP/TLS connection, parses the request within its configured inspection limits, runs admission and endpoint selection, and manages the backend connection. Engine-specific P/D orchestration also runs in this serving path.
Why AI routing needs the userspace proxy
L4 eBPF forwarding never sees HTTP — it makes its decision from the packet's 5-tuple before
any request body arrives. Model-name routing, KV-cache-aware prefix matching, P/D request
splitting, and SSE stream handling all require reading the HTTP request (and sometimes the
body). That inspection happens in the userspace fullproxy, which is why mode: 4 is the
prerequisite for every AI feature. See Running Modes.
Where AI routing hooks in¶
Inference-aware behavior is layered onto the fullproxy request path. For a request on an AI-enabled rule, the gateway proceeds roughly as follows:
- Terminate and parse. The fullproxy accepts the client connection (optionally
terminating TLS per the rule's
securitymode) and reads the HTTP request. - Model selection. If
model_namepools are configured, the requested model (from theX-Modelheader or the bodymodelfield) picks the endpoint pool;""is the catch-all. - Endpoint selection. The rule's
selalgorithm chooses an endpoint within the pool. For LLM fleets this is typically CHWBL prefix affinity (sel: 8/10) or, when a KV-cache event stream is wired, engine-exact KV routing that places the request on the endpoint already holding the longest matching prompt prefix. - P/D orchestration (optional). With
pd_disagg_modeenabled, the proxy applies the selected engine dialect: sequential vLLM prefill/decode, concurrent SGLang prefill/decode, or sequential TensorRT-LLM context/generation. llama.cpp P/D is rejected at rule creation. - Stream relay. With
sse_modeenabled, the response is relayed as SSE with idle-timeout suppression while the stream is active, a wall-clock cap, and optional backend keepalive.
The KV inventory services run alongside this path. As supported backends publish events, the gateway updates its per-endpoint hash inventory so KV-exact selection can reflect observed cache locality. Hash inventories contain compact hashes, not KV tensors.
Coexistence with classic load balancing¶
Because AI fields are opt-in per rule, an inference-gateway node can host AI rules and classic
L4 rules side by side. A vLLM VIP (mode: 4, KV-aware) and a plain TCP service-type load
balancer can live on the same gateway; the classic rule uses the eBPF fast path. AI runtime
state such as KV inventories and circuit-breaker state has separate failover limitations; do
not infer stateful AI high availability from L4 cluster support alone. See
HA Limitations.
Next¶
- Running Modes — the
modeenum and whymode: 4is the AI prerequisite. - LB Algorithms — the
selselection policies, including CHWBL. - AI Gateway Overview — the full inference feature set.
- Quickstart — a working rule end to end.