General Compute
General Compute offers dedicated inference hardware under contract. The company owns the racks and handles site selection, bring-up and operations; customers receive dedicated capacity and root access to their machines. Its offer targets teams deploying models at a defined traffic level rather than users buying a small bundle of chatbot requests.
Specify the workload before choosing the hardware
General Compute describes splitting model serving between GPU prefill and purpose-built decode chips. Its listed hardware families include NVIDIA, SambaNova, Cerebras, Positron and d-Matrix. The proposed access surface can be bare metal or a combined endpoint, with one contract and a set of service-level agreements.
The company's pricing page asks about prompt length, output length, streaming and concurrency. These are useful inputs for a capacity discussion because a short interactive response and a long batch generation place different demands on a deployment. Provide a representative workload rather than selecting hardware from a headline token rate.
Model ownership also affects the offer. General Compute distinguishes hosted open models from private weights and says commitment length influences reserved-rack pricing. The inspected pricing page provides a request form, not a public per-token tariff. Obtain a quote that states the model, capacity and contract term before comparing it with another service.
Read the benchmark conditions alongside the speed claim
The homepage presents coding-session comparisons and an interactivity chart. It labels the chart as an internal benchmark set, with a SambaNova configuration and an NVIDIA comparison approximated from SemiAnalysis InferenceX. These are company-presented measurements under stated conditions, not Franklin's independent performance test.
A high displayed token rate does not establish your application's response time under its own concurrency and prompt lengths. Ask for evidence at the operating point you intend to purchase, including how the deployment handles your weights and serving requirements. Keep per-user rate separate from total rack throughput when reviewing the offer.
Confirm the deployment slot and availability
The site announces a multi-year Cerebras agreement with deployment live in Q1 2027. That future date should not be collapsed into a claim that the announced capacity is available today. The page also discusses different chip families reaching production or volume at different stages. Ask which hardware is available for your requested slot.
General Compute says it handles power and cooling in existing US colocation facilities and can site capacity alongside a customer's fleet where possible. A purchase conversation should make the location, operational responsibilities and service commitments explicit. The hardware catalogue is a starting point; the signed deployment terms determine what capacity you receive.
