Back to AI Research

AI Research

You Cannot Pick a Provider From the Price List: Mar... | AI Research

Key Takeaways

  • You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference Choosing an LLM for a request is only part of the deployme...
  • Existing LLM routers choose among models using static per-model costs.
  • We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it.
  • Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list.
  • We formulate same-model provider selection as a price-taker market-aware routing problem.
Paper AbstractExpand

Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.

You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference

Choosing an LLM for a request is only part of the deployment decision. When open-weight models are hosted by multiple competing providers, the same model can produce different results depending on who serves it. This paper studies that overlooked provider-selection problem and proposes routing methods that balance cost, quality, latency, and reliability. Its central claim is that the paper’s study of market-aware routing for open-weight LLM inference shows why a provider cannot be selected from a price list alone.

What the paper measures

The authors examine six open models across multiple providers, task types, and three measurement waves. They pin requests to individual providers rather than allowing automatic fallback, then record accuracy, realized cost, latency, and availability. The tasks include math, extraction, sentiment classification, code-output prediction, GSM8K, MMLU, and HumanEval.
The measurements reveal substantial variation among providers serving the same model. Prices can differ by as much as 8.8 times, while latency can differ by 68 times. Accuracy may also vary by tens of percentage points on difficult tasks. In five measured model–task cells, the cheapest provider fell below the paper’s quality threshold.
Price does contain one useful signal: speed. Across 21 benchmark-scale cells, higher-priced providers were consistently faster, with a median price–latency Spearman correlation of −0.61. But price did not reliably predict answer quality or availability. The corresponding median correlations were +0.05 for accuracy and 0.00 for availability. In other words, paying more often bought lower latency, but not necessarily better answers or a safer endpoint.
The authors also find that provider quality is task-dependent. A deployment can perform normally on knowledge-oriented tasks, degrade on code, and fail severely on multi-step reasoning. This makes a simple provider-wide label such as “safe” or “unsafe” inadequate. A useful routing map must be indexed by model, provider, and task.

How the routing approach works

The first policy is a measured-map router. For each model–task combination, it chooses the cheapest provider that is both healthy and close enough in measured quality to the best available provider. In the primary setting, “close enough” means within five percentage points of the best measured accuracy, while availability must exceed 90%.
This formulation treats provider selection as a constrained cost-minimization problem. The router seeks low cost while requiring a task-conditioned quality floor, sufficient availability, and, when relevant, an acceptable latency. The approach is separate from ordinary model routing: an upstream system may first choose an 8B or 70B model, after which this provider layer decides who should serve that model.
The measured map is useful but temporary. Prices may remain stable while the best route changes because of provider exit, rate limiting, recovering availability, or silent quality degradation. Across the measurement waves, the selected route changed in several comparable cells even without major price movement. The paper therefore introduces FACET, an online provider router that treats each provider–task pair as a separate feasibility “facet.”
FACET does not immediately send traffic to a newly attractive cheap provider. Instead, it probes that provider while routing requests to a trusted anchor. Once enough labelled observations show that the candidate is sufficiently close to the best same-task provider and above the quality floor, FACET certifies it for service. After certification, it continues monitoring the endpoint and quarantines it if performance deteriorates. Unknown or uncertified facets fail safe to the anchor.

Results that stand out

The measured-map policy achieves 94% mean accuracy at 1.57 times the cost of always choosing the cheapest provider, while maintaining a zero below-floor rate in the reported in-sample comparison. The single-best provider oracle reaches 95% accuracy, but the premium policy costs 3.67 times the cheapest-provider cost and still sends 17% of cells below the quality floor. Pure cheapest-provider routing costs 1.0 times the baseline but has a 28% below-floor rate.
The savings remain substantial when accounting for statistical uncertainty, although they shrink. On a 14-cell benchmark subset, requiring a one-sided 95% confidence interval reduces the median saving from 55% to 46%. The authors attribute much of the supported saving to a mismatch between price and quality: the most expensive provider is often not the most accurate one.
Provider-aware routing also remains useful beneath an upstream model router. In a combined evaluation, cheapest-provider routing sent 55.9% of MMLU queries and 57.5% of HumanEval queries through below-floor endpoints. Provider-aware routing reduced both rates to zero, with only modest increases in total cost in the reported experiment.
Latency creates an explicit trade-off. Among cells with multiple quality-feasible providers, the cheapest feasible endpoint had a median p95-latency penalty of 7.8 times compared with the fastest feasible endpoint. A 10-second latency requirement preserved all 12 evaluated cells at 1.19 times unconstrained routing cost, while a 1-second requirement left only six feasible cells and raised cost to 2.62 times the unconstrained policy.

What to keep in mind

The paper’s results depend on measurement choices, task assignment, quality thresholds, and feedback quality. The authors test imperfect task assignment and sparse feedback, but systematic evaluator bias creates a trust boundary: if the quality signal is wrong in a consistent way, certification can admit a bad provider. They identify ground-truth probes and audits as safeguards.
The approach also requires ongoing probing and monitoring. A snapshot of provider behavior can become stale, and task labels may be unavailable or uncertain. FACET addresses these issues through conservative certification and fallback, but that safety comes with measurement overhead and potentially greater reliance on the anchor provider.
The broader lesson is operational rather than merely economic. Open-weight models are not stable services independent of their hosts. As SWE-Serve’s focus on complete, live inference systems also reflects, production inference depends on serving behavior as well as model capability. For routing systems, choosing the model is therefore only the first decision; choosing a provider requires direct, task-specific, and continuously updated evidence.

Comments (0)

No comments yet

Be the first to share your thoughts!