You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
Choosing an LLM for a request is only part of the deployment decision. When open-weight models are hosted by multiple competing providers, the same model can produce different results depending on who serves it. This paper studies that overlooked provider-selection problem and proposes routing methods that balance cost, quality, latency, and reliability. Its central claim is that the paper’s study of market-aware routing for open-weight LLM inference shows why a provider cannot be selected from a price list alone.
What the paper measures
The authors examine six open models across multiple providers, task types, and three measurement waves. They pin requests to individual providers rather than allowing automatic fallback, then record accuracy, realized cost, latency, and availability. The tasks include math, extraction, sentiment classification, code-output prediction, GSM8K, MMLU, and HumanEval.
The measurements reveal substantial variation among providers serving the same model. Prices can differ by as much as 8.8 times, while latency can differ by 68 times. Accuracy may also vary by tens of percentage points on difficult tasks. In five measured model–task cells, the cheapest provider fell below the paper’s quality threshold.
Price does contain one useful signal: speed. Across 21 benchmark-scale cells, higher-priced providers were consistently faster, with a median price–latency Spearman correlation of −0.61. But price did not reliably predict answer quality or availability. The corresponding median correlations were +0.05 for accuracy and 0.00 for availability. In other words, paying more often bought lower latency, but not necessarily better answers or a safer endpoint.
The authors also find that provider quality is task-dependent. A deployment can perform normally on knowledge-oriented tasks, degrade on code, and fail severely on multi-step reasoning. This makes a simple provider-wide label such as “safe” or “unsafe” inadequate. A useful routing map must be indexed by model, provider, and task.
How the routing approach works
The first policy is a measured-map router. For each model–task combination, it chooses the cheapest provider that is both healthy and close enough in measured quality to the best available provider. In the primary setting, “close enough” means within five percentage points of the best measured accuracy, while availability must exceed 90%.
This formulation treats provider selection as a constrained cost-minimization problem. The router seeks low cost while requiring a task-conditioned quality floor, sufficient availability, and, when relevant, an acceptable latency. The approach is separate from ordinary model routing: an upstream system may first choose an 8B or 70B model, after which this provider layer decides who should serve that model.
The measured map is useful but temporary. Prices may remain stable while the best route changes because of provider exit, rate limiting, recovering availability, or silent quality degradation. Across the measurement waves, the selected route changed in several comparable cells even without major price movement. The paper therefore introduces FACET, an online provider router that treats each provider–task pair as a separate feasibility “facet.”
FACET does not immediately send traffic to a newly attractive cheap provider. Instead, it probes that provider while routing requests to a trusted anchor. Once enough labelled observations show that the candidate is sufficiently close to the best same-task provider and above the quality floor, FACET certifies it for service. After certification, it continues monitoring the endpoint and quarantines it if performance deteriorates. Unknown or uncertified facets fail safe to the anchor.
Results that stand out
The measured-map policy achieves 94% mean accuracy at 1.57 times the cost of always choosing the cheapest provider, while maintaining a zero below-floor rate in the reported in-sample comparison. The single-best provider oracle reaches 95% accuracy, but the premium policy costs 3.67 times the cheapest-provider cost and still sends 17% of cells below the quality floor. Pure cheapest-provider routing costs 1.0 times the baseline but has a 28% below-floor rate.
The savings remain substantial when accounting for statistical uncertainty, although they shrink. On a 14-cell benchmark subset, requiring a one-sided 95% confidence interval reduces the median saving from 55% to 46%. The authors attribute much of the supported saving to a mismatch between price and quality: the most expensive provider is often not the most accurate one.
Provider-aware routing also remains useful beneath an upstream model router. In a combined evaluation, cheapest-provider routing sent 55.9% of MMLU queries and 57.5% of HumanEval queries through below-floor endpoints. Provider-aware routing reduced both rates to zero, with only modest increases in total cost in the reported experiment.
Latency creates an explicit trade-off. Among cells with multiple quality-feasible providers, the cheapest feasible endpoint had a median p95-latency penalty of 7.8 times compared with the fastest feasible endpoint. A 10-second latency requirement preserved all 12 evaluated cells at 1.19 times unconstrained routing cost, while a 1-second requirement left only six feasible cells and raised cost to 2.62 times the unconstrained policy.
What to keep in mind
The paper’s results depend on measurement choices, task assignment, quality thresholds, and feedback quality. The authors test imperfect task assignment and sparse feedback, but systematic evaluator bias creates a trust boundary: if the quality signal is wrong in a consistent way, certification can admit a bad provider. They identify ground-truth probes and audits as safeguards.
The approach also requires ongoing probing and monitoring. A snapshot of provider behavior can become stale, and task labels may be unavailable or uncertain. FACET addresses these issues through conservative certification and fallback, but that safety comes with measurement overhead and potentially greater reliance on the anchor provider.
The broader lesson is operational rather than merely economic. Open-weight models are not stable services independent of their hosts. As SWE-Serve’s focus on complete, live inference systems also reflects, production inference depends on serving behavior as well as model capability. For routing systems, choosing the model is therefore only the first decision; choosing a provider requires direct, task-specific, and continuously updated evidence.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!