Skip to main content
50% off all plans, limited time. Starting at $2.48/mo
16 min left
AI & Machine Learning

Top AI Chip Companies and Manufacturers: The NVIDIA Alternatives That Actually Ship

J By Jeremy 16 min read
Top AI chip companies: a graphics card, an accelerator module, a silicon wafer and server boards arranged on a dark red background

You can order a Tenstorrent Blackhole developer card this week. It costs $999, it is in stock, and it ships in about two weeks. You cannot buy or rent a Meta MTIA through a public channel; Meta describes MTIA as silicon it deploys for its own workloads.

Both companies get counted among the top AI chip companies. The distance between those two situations is what this article is about.

What follows is vendor by vendor: what you can get, through which channel, and what the software costs you once you have it.

TL;DR

  • Broadly rentable: AMD Instinct still has the broadest non-NVIDIA public-cloud footprint, but the product line has moved on. AMD launched the MI400 Series in July 2026, while MI300X, MI325X, and MI350-series systems remain the ones you are most likely to find for rent.
  • Single-cloud: Google TPUs stay on Google Cloud, with TPU7x Ironwood generally available since March 2026. AWS Trainium and Inferentia stay on AWS.
  • API plus dedicated infrastructure: Groq, Cerebras, and SambaNova all offer hosted inference, but API-only is no longer accurate. Each now has a dedicated or on-premises path as well.
  • Hardware you can procure: Tenstorrent sells Blackhole cards directly. Qualcomm presents its Dragonfly data-center accelerators through an enterprise contact-sales path rather than as roadmap announcements alone.
  • Limited or internal: Intel Gaudi 3 remains available through select IBM Cloud deployments while Crescent Island is expected to enter customer sampling in the second half of 2026. Etched says its first racks ship to contract customers in summer 2026. Meta MTIA and Microsoft Maia remain internal or Azure-integrated rather than public self-serve accelerators.

What This Article Doesn't Cover

Three things sit deliberately outside the frame, and every vendor status below reflects August 2026.

  • Pricing. The question here is whether you can get the hardware, not what an hour on it costs.
  • Forecasting. No calls about who wins.
  • Consumer and desktop cards. The subject is data-center and cloud accelerators.

The Five Ways You Can Get Compute on an AI Accelerator

Five access tiers for AI accelerators in August 2026, running from most open to most restricted: many clouds for AMD Instinct, a single cloud for Google TPU7x and AWS Trainium, hosted plus dedicated infrastructure for Groq, Cerebras and SambaNova, hardware or enterprise sales for Tenstorrent, Qualcomm and Etched, and internal only for Meta MTIA and Microsoft Maia 200

Five access paths cover the accelerators in this article: rental across multiple clouds, rental through a single cloud, hosted API access, dedicated or directly obtainable hardware, and limited enterprise access for products that are not broadly self-service. Some companies now span more than one of these paths. A separate group remains internal-only, with no public way to rent or buy the silicon.

Access is a status, not a permanent property of the company. It changes as new generations launch, clouds add or drop hardware, and vendors open or close procurement channels. It also says nothing about quality or scale. It describes how you can get the compute, and that channel decides whether a name is something you can act on.

NVIDIA holds a large majority of data-center AI accelerator revenue, it is on virtually every cloud, and CUDA is the surface everything else gets measured against. The other AI accelerator companies here are interesting because of how differently they are reached.

Vendor and productAccess tierHow you get itSoftware stackWorkloadStatus
NVIDIA (baseline, not an alternative)Many cloudsNearly every cloud providerCUDATraining and inferenceGenerally available
AMD Instinct MI300XMany cloudsVultr, TensorWave, Oracle Cloud, DigitalOcean, Crusoe, Hot Aisle, RunPod, Seeweb, CirrascaleROCmTraining and inferenceGenerally available
AMD Instinct MI325XMany cloudsVultr, TensorWave, DigitalOceanROCmTraining and inferenceGenerally available
Google TPU7x (Ironwood)One cloudGoogle CloudJAX, and PyTorch via PyTorch/XLATraining and inferenceGenerally available since March 31, 2026
AWS TrainiumOne cloudAWS EC2 Trainium instances and UltraServers (Trn2, Trainium3)Neuron SDKTraining and inferenceGenerally available
AWS Inferentia2One cloudAWS EC2 Inf2 instancesNeuron SDKInferenceGenerally available
Groq LPUHosted API and dedicated infrastructureGroqCloud and GroqMetalOpenAI-compatible API and the Groq stackInferenceAvailable
Cerebras WSE-3 and CS-3Hosted API, enterprise cloud, on-premisesCerebras Cloud and CS-3 systemsCerebras software stackTraining and inferenceAvailable
SambaNova SN50Hosted API, dedicated, on-premisesSambaCloud and SambaRackSambaStackInferenceCustomer shipping begins in the second half of 2026
Intel Gaudi 3Select cloudIBM CloudIntel Gaudi software suiteTraining, fine-tuning and inferenceSelect availability
Tenstorrent BlackholeBuy the hardwareDirect from TenstorrentOpen-source TT-Metalium stackInference and dev workIn stock, ships in about two weeks
Qualcomm AI200, AI250 and AI300Enterprise salesContact salesQualcomm AI software stackInferenceAI200 expected in 2026, AI250 sampling expected in 2027, AI300 sampling expected in 2028
Etched SohuContract deploymentsEnterprise contractsEtched software stackTransformer inferenceFirst racks shipping in summer 2026
Meta MTIANot obtainableNo external pathInternal Meta toolingMeta's own ranking, recommendation and GenAI workInternal only
Microsoft Maia 200Internal, Azure-integratedMicrosoft datacentersMaia SDKInferenceDeployed internally, with no public self-serve instance

AMD Instinct Is the One You Can Rent Almost Anywhere

AMD Instinct has the widest provider footprint of the NVIDIA alternatives here, with MI300X and MI325X available across multiple cloud and neocloud providers rather than being tied to one hyperscaler. AMD launched the MI400 Series in July 2026, led by the MI455X, but the older generations still matter because they are the ones you are more likely to find in existing rental deployments. ROCm remains the shortest software port in this article for most existing PyTorch workloads.

The hardware is not the constraint. MI300X carries 192GB of HBM3 at 5.3 TB/s across 304 CDNA3 compute units in a 750W module. MI325X pushes that to 256GB of HBM3E at 6 TB/s and 1000W, a capacity that landed lower than the 288GB AMD announced at first.

CUDA cores are what NVIDIA counts, where AMD counts compute units, and the two do not convert, which is why memory and bandwidth compare more usefully across vendors.

The software situation has changed, and the reputation lags it. ROCm support for PyTorch is upstreamed into the official PyTorch repository rather than living in a community fork, and current ROCm releases track current frameworks. ROCm 7.14.0 supports PyTorch 2.12, alongside AMD's current integrations for the wider AI software ecosystem.

My read is that the friction people still describe is not throughput and not framework support. It is archaeology. When a CUDA setup breaks at three in the morning, someone has already hit that error string and written down the fix. On ROCm the search is thinner, so the time goes into looking rather than fixing. That cost never shows up in a benchmark.

AMD's accelerator work also reaches past the data center, which matters if you want cheap hardware to learn the stack on.

AMD's mini PC cluster demo runs on the same company's consumer silicon.

Google TPU and AWS Trainium Are Rentable on Exactly One Cloud Each

Google's TPU7x and AWS's Trainium are both generally available and both locked to a single provider. TPU runs on Google Cloud. Trainium and Inferentia run on AWS. Neither gives you the multi-provider portability of AMD or NVIDIA, and that constraint can matter more than any individual specification.

TPU7x is the first release in Google's seventh-generation Ironwood family and became generally available on March 31, 2026. Google positions it for large-scale training and inference. Trillium v6e remains available, while Ironwood remains Google's current generally available TPU generation. Google has also announced eighth-generation TPU 8t and TPU 8i, but both are still listed as coming soon.

The most useful fact about TPUs is not on the spec sheet. There is a free on-ramp: Colab and Kaggle both offer TPU runtimes, and the TPU Research Cloud grants time to researchers. That free on-ramp gives you a way to test the TPU software path before committing cloud budget, although the free services should not be assumed to provide the latest Ironwood hardware specifically.

AWS reaches its own silicon through Trainium instances and UltraServers, plus Inf2 instances for Inferentia2. Trainium3 UltraServers reached general availability in December 2025 and scale to 144 chips, up from 64 on the previous Trn2 UltraServers. Everything routes through the Neuron SDK. AWS frames the porting cost more optimistically than most vendors do.

The AWS Trainium page states that vLLM, HuggingFace Transformers, and TorchTitan run natively on Trainium with no custom code and no porting effort. That is the vendor's claim about its own product, not an independently verified result.

Both sit inside a trade worth making deliberately. You gain a second silicon supplier and lose the ability to move that workload to another cloud without redoing the work, which costs almost nothing to a team already committed to one provider and everything to a team built around avoiding commitment.

What has to fit in memory constrains the choice first: 144GB per Trainium3 chip is a different planning problem from Ironwood's 192GB, whoever sells it.

Groq, Cerebras, and SambaNova Go Beyond Hosted APIs

Groq, Cerebras, and SambaNova all offer hosted inference, but the API is no longer the whole access story. Groq pairs GroqCloud with dedicated GroqMetal infrastructure. Cerebras offers cloud access alongside CS-3 systems that can run on premises. SambaNova offers SambaCloud alongside SambaRack and SambaStack. The distinction now is how much infrastructure control you need and how far into enterprise procurement you are willing to go.

Groq's LPU packs 230MB of on-die SRAM with 80 TB/s of internal bandwidth, a design that trades memory capacity for latency, and GroqCloud exposes it as an OpenAI-compatible endpoint. The ownership story attached to it carries dates worth keeping. On 24 December 2025, Groq entered a non-exclusive inference technology licensing agreement with NVIDIA, and said founder Jonathan Ross and president Sunny Madra would join NVIDIA. Groq continues to operate independently. Simon Edwards, previously its chief financial officer, stepped up to lead it at the time of the deal; check the company for current leadership before citing it.

Groq's own announcement says that GroqCloud will continue to operate without interruption.

At GTC in March 2026, NVIDIA unveiled the NVIDIA Groq 3 LPU inside its Vera Rubin platform.

NVIDIA's Vera Rubin announcement describes an LPX rack of 256 LPU processors, available in the second half of 2026.

Cerebras takes the opposite physical approach: one wafer-scale part with 900,000 cores and 44GB of on-chip SRAM, rated at 125 PFLOPs.

Cerebras's inference page offers $5 in free credit and describes the migration as just two code changes. One distinction still matters for planning: the self-serve API is the easy inference path, while training, fine-tuning, dedicated capacity, and on-premises CS-3 deployments sit farther up the enterprise path.

SambaNova's current hardware story centers on SambaRack SN50, built around 16 fifth-generation SN50 RDUs per rack. SambaCloud remains the hosted path, while SambaRack and SambaStack provide dedicated infrastructure that can run on premises or in hosted environments.

Hosted access is still the easiest entry point for all three, but it no longer describes the full product.

Tenstorrent Sells Hardware You Can Own

Tenstorrent is the one vendor in this article whose accelerator you can buy outright as an individual, at a price that is not a capital decision.

Tenstorrent's hardware page lists the Blackhole p100a at $999 and the p150a and p150b at $1,399, in stock and shipping in about two weeks. The stack is fully open source.

The cards carry 120 Tensix cores and 16 RISC-V cores, with 28GB of GDDR6 on the p100a and 32GB on the p150 pair, drawing up to 300W. The p150a and p150b add QSFP-DD Ethernet ports the p100a lacks, which is what you buy if you intend to link cards together. HPCwire reported that the rack-scale Galaxy Blackhole reached general availability on 28 April 2026 from $110,000, and the older Wormhole line remains on sale, including a $12,000 TT-LoudBox workstation.

The caveat is as decision-relevant as the price, and it points at that older card. Most of Tenstorrent's documentation, tutorials, and verified model support still target Wormhole. Blackhole's software support is earlier in its cycle. Concretely: this suits someone willing to do genuine porting work and read source when the docs run out. It does not suit someone who wants existing PyTorch code to run unmodified, which is a perfectly reasonable thing to want.

Qualcomm, Intel, and Etched Have Narrower Access Paths

These three sit behind narrower access paths for unrelated reasons. Qualcomm is building out a multi-generation data-center accelerator roadmap around enterprise deployments. Intel Gaudi 3 remains available through select IBM Cloud deployments while Crescent Island moves toward customer sampling. Etched says its first Sohu racks ship in summer 2026 against enterprise contracts rather than through a self-serve cloud service.

Qualcomm's data-center roadmap now spans Dragonfly AI200, AI250, and the newly announced AI300, with an annual accelerator cadence focused on inference. This remains an enterprise-oriented path rather than the kind of self-serve accelerator rental you can add to a cloud account today.

Intel is the instructive entry because Gaudi 3 never achieved a broad cloud footprint, but it has not disappeared. IBM Cloud still offers Gaudi 3 accelerated virtual-server profiles under select availability, including deployments in Dallas, Washington DC, and Frankfurt. Intel's next step is Crescent Island, an inference-focused GPU expected to enter customer sampling in the second half of 2026. Gaudi 3 is a narrow option today, not a concluded one.

Etched is building Sohu, an ASIC specialized for transformer inference. The company has raised $800 million, has working A0 silicon, reports more than $1 billion in signed customer contracts, and first racks are scheduled to ship in summer 2026. Its performance claims are its own and have not been independently benchmarked, so the throughput numbers are not the useful part. The software stack is the useful part. Etched is building the chip, rack, and software environment as one specialized platform, so you should not assume the CUDA-era serving stack transfers unchanged. That can work for a contract customer running transformer inference at enormous scale, but it is a much larger commitment for a team that depends on accelerator portability.

Internal-Only Silicon and the Firms That Build Chips for Others

Two situations look like an AI chip company from the outside and behave nothing alike. Meta and Microsoft design accelerators they operate themselves and sell to nobody. Broadcom and Marvell design and implement silicon for other companies' architectures. Neither group has a product a developer can rent, for reasons that have nothing in common.

Meta's MTIA is the cleanest case of the first kind.

Meta's own engineering post describes deploying hundreds of thousands of MTIA chips for inference workloads across both organic content and ads on its apps, with MTIA 300 already in production for ranking and recommendation training and MTIA 400, 450, and 500 planned on a cadence of roughly six months. That is an enormous amount of silicon you will never rent an hour of.

Microsoft's case is softer at the edge. Maia 200, introduced in January 2026, is Microsoft's current in-house inference accelerator and is already deployed in its datacenters, with the Maia SDK available in preview. Microsoft does not currently document a public self-serve Maia instance, so it still belongs on the internal or Azure-integrated side of this comparison.

Broadcom and Marvell are a different thing again, worth clearing up because their names appear constantly next to NVIDIA's. Google's TPU is a useful example. Google owns the TPU product and architecture, while Broadcom contributes custom-silicon design and implementation work that can include interconnects, packaging, and the path from architecture to manufacturable silicon. Broadcom does not sell a TPU, or another general-purpose accelerator, directly to developers. Its AI-segment revenue, reported from its fiscal Q1 2026 earnings, reached $8.4 billion, up 106 percent year over year, which explains why the name is everywhere. It still is not a chip you can rent.

Which is where the word manufacturers starts carrying weight it cannot hold. Broadcom and Marvell do not manufacture chips, and neither do most of the AI chip manufacturers listed above: nearly all of them are fabless, meaning they design silicon and contract the fabrication out to someone else entirely.

Matching an Access Path to Your Workload

Training at scale narrows the practical choices quickly: broadly available AMD Instinct, single-cloud TPU or Trainium, or an enterprise Cerebras deployment. Production inference opens more paths because Groq, Cerebras, and SambaNova all provide hosted services while also offering some form of dedicated infrastructure. Experimentation is wider again, ranging from free TPU access to a Tenstorrent card you can put in your own machine.

Those boundaries are softer than a table can show. The question is not only what the silicon can do but where you can reach it, how much of your existing software stack survives the move, and whether the resulting workload can move somewhere else later. Cerebras can train and infer across enterprise infrastructure, while Trainium remains tied to AWS.

None of this dislodges NVIDIA. For most teams the live question is not whether to leave CUDA but whether a second stack is worth maintaining alongside it.

Which card to size for is the remaining decision if you stay where you are, and that is a different exercise from this one.

The channel decides more than the chip does. A card on your desk and the same class of silicon behind somebody else's API are not the same tool even when the throughput lands close, because one lets you try things you have not thought of yet and the other bills you per token for what you already know how to ask.

Frequently Asked Questions

Can You Actually Rent a Google TPU?

Yes. TPU7x, Google's current generally available Ironwood TPU, is available on Google Cloud and only there. Older TPU generations are also available, and Colab, Kaggle, and the TPU Research Cloud provide free TPU access for some workloads. JAX remains the native path, while PyTorch runs through PyTorch/XLA.

Which NVIDIA Alternative Runs Existing PyTorch Code With the Least Work?

AMD Instinct, in most cases: ROCm is an officially supported PyTorch backend, so the code path already exists. Google's TPU runs PyTorch through the PyTorch/XLA bridge, a translation layer. AWS states that vLLM, HuggingFace Transformers, and TorchTitan run on Trainium with no porting effort, which is the vendor's own claim. At the far end, Etched's Sohu uses its own specialized software stack, so you should expect substantially more porting work than with ROCm or PyTorch/XLA.

Why Are Broadcom and Qualcomm Called NVIDIA Competitors If You Cannot Rent Their Chips?

For different reasons. Broadcom works on custom silicon for companies building their own accelerators rather than selling a standard accelerator directly to developers. Qualcomm designs its own Dragonfly accelerators, with AI200, AI250, and AI300 now on its data-center roadmap, but access is enterprise-oriented rather than a normal self-serve cloud rental. Both compete in AI infrastructure without giving developers the same public access model as NVIDIA or AMD.

Can You Buy an AI Accelerator Outright Instead of Renting One?

Yes. Tenstorrent sells Blackhole developer cards directly, from $999 for the p100a to $1,399 for the p150a and p150b, in stock and shipping in about two weeks, with a fully open-source software stack. The caveat matters as much as the price: most documentation and verified model support still target the older Wormhole card, so budget time for porting rather than expecting existing code to run unchanged.

Share

Discussion

Comments

Sign in to join the discussion.

More from the blog

Keep reading.

Ready to deploy? From $2.48/mo.

Independent cloud, since 2008. AMD EPYC, NVMe, 40 Gbps. 14-day money-back.