RISC-V vs NVIDIA for AI Inference:
A Total Cost of Ownership Analysis
Why open-ISA silicon may reshape the economics of production AI
Executive Summary
The AI inference market is dominated by NVIDIA GPUs, which command an estimated 80–95% share of data-centre AI accelerator deployments globally.1 This dominance has driven GPU prices to historic highs and created structural supply constraints that affect every organisation deploying production AI workloads.
A new class of alternatives is emerging. Tenstorrent, led by legendary chip architect Jim Keller, has shipped the Galaxy server platform — a RISC-V-based AI inference system that takes a fundamentally different approach to silicon economics. Rather than competing on raw FP16 training throughput, Galaxy targets the production inference workloads that constitute the majority of real-world AI compute demand.
This analysis constructs a five-year total cost of ownership (TCO) model comparing a 100-server Tenstorrent Galaxy cluster against an equivalent NVIDIA B200 deployment for production LLM inference. We examine capital costs, energy consumption, cooling infrastructure, software licensing, and the less-quantifiable but increasingly material costs of vendor dependency and export-control exposure.
For April 2026 general-availability list prices ($110,000 entry 6U / 32-chip; four-Galaxy supercluster from $440,000) and the DGX 3–5× trade-off including TT-Metal porting cost, see the hardware hub Tenstorrent Galaxy Blackhole price. This research page is TCO methodology — not a live campus hall.
For sovereignty-sensitive inference workloads — where vendor independence, data jurisdiction, and supply-chain resilience are weighted alongside raw performance — RISC-V-based deployments present a compelling TCO trajectory. The gap widens as energy costs rise and RISC-V software toolchains mature. However, NVIDIA retains decisive advantages in training workloads and ecosystem breadth. This is not a wholesale replacement thesis; it is a workload-specific optimisation argument.
The Hardware: Two Architectures, Two Philosophies
Before building a cost model, it is essential to understand what we are comparing. These are not interchangeable commodity servers — they represent distinct architectural philosophies with different strengths and operational characteristics.
Tenstorrent Galaxy
Tenstorrent's Galaxy platform is built around the Blackhole RISC-V AI accelerator. Each Galaxy server contains 32 Blackhole chips in a 6U air-cooled chassis drawing 8–10 kW operating (not nameplate; ~12 kW max).2 The system uses the open RISC-V instruction set architecture, meaning the silicon design is auditable down to the ISA level — a property with significant implications for data-sovereignty use cases.
Galaxy is positioned explicitly for inference workloads. Tenstorrent has published benchmarks showing competitive tokens-per-second throughput on large language models, though independent third-party benchmarks at scale remain limited as of mid-2026. The system ships with Tenstorrent's open-source TT-Metalium software stack.
NVIDIA B200
The NVIDIA B200 (Blackwell architecture) is the current-generation flagship AI accelerator. Each B200 GPU has a rated TDP of approximately 1,000 W and requires liquid cooling for sustained operation at rated performance.3NVIDIA's CUDA ecosystem provides the industry's most mature AI software stack, with deep framework integration across PyTorch, TensorRT, and the full NVIDIA AI Enterprise suite.
B200 excels at both training and inference workloads, with particularly strong FP8/FP4 mixed-precision performance. However, it carries the cost profile of a general-purpose accelerator designed to dominate every AI workload category simultaneously.
Galaxy is a purpose-built inference platform with an air-cooled thermal envelope. B200 is a general-purpose AI accelerator requiring liquid cooling infrastructure. This distinction cascades through every line item of the TCO model — from capital infrastructure to ongoing energy costs.
TCO Model Methodology
Our five-year TCO framework evaluates costs across six categories. Where data is publicly available, we use verified figures. Where AGICY-specific projections are required — particularly for deployment-specific configurations — these are clearly labelled.
1. Capital Cost (Silicon + Servers + Networking)
Hardware acquisition is the largest single line item in any AI infrastructure deployment. NVIDIA GPU pricing varies significantly by customer, volume, and contract structure — published list prices are rarely what large buyers actually pay. Tenstorrent Galaxy pricing is available through direct engagement but has not been publicly disclosed at the per-unit level as of mid-2026.
For this analysis, we use analyst consensus ranges for NVIDIA B200 pricing ($30,000–$40,000 per GPU) and Tenstorrent's publicly indicated price positioning, which targets significant per-unit cost reduction relative to NVIDIA equivalents.4
2. Energy Cost
Energy cost is computed as: TDP × utilisation factor × hours × cost per kWh. We assume 85% average utilisation (standard for production inference clusters), 8,760 hours per year, and a blended electricity rate of €0.10/kWh — consistent with Cyprus industrial power rates under long-term contract.5
3. Cooling Infrastructure
This is where architectural differences create compounding cost divergence. Air-cooled systems (Galaxy) operate within standard data-centre environments with a power usage effectiveness (PUE) of approximately 1.3–1.4. Liquid-cooled systems (B200) require purpose-built cooling distribution units, rear-door heat exchangers or direct-to-chip cooling loops, and associated plumbing — driving PUE as low as 1.1 but with substantially higher capital expenditure and maintenance overhead.6
4. Software Licensing
NVIDIA AI Enterprise licensing costs $4,500 per GPU per year for the full software stack.7 This is optional — organisations can use open-source CUDA toolkits — but enterprise customers deploying at scale typically require the support, security patches, and certified containers that the enterprise license provides.
Tenstorrent's TT-Metalium stack is fully open-source under the Apache 2.0 license, with no per-unit software fees. The trade-off is a less mature ecosystem with fewer pre-optimised model implementations.
5. Vendor Dependency Risk
This category is difficult to quantify in dollar terms but increasingly material. NVIDIA's market dominance means customers face single-vendor pricing power, supply allocation decisions outside their control, and exposure to US export-control policy changes. The October 2022, October 2023, and January 2025 BIS rule updates have repeatedly restricted which customers can purchase which NVIDIA products — creating procurement uncertainty even for allied-nation buyers.8
6. Maintenance & Operations
Ongoing operational costs including hardware warranties, spare parts, staffing, and facility maintenance. We apply industry-standard estimates of 8–12% of capital cost annually for both platforms.
Five-Year TCO Comparison: 100-Server Cluster
The following table presents our modelled TCO for a 100-server inference cluster under each architecture. All figures are for the cluster as a whole over five years. AGICY-specific projections are marked; all other figures derive from public sources.
| Cost Category | Tenstorrent Galaxy (RISC-V) | NVIDIA B200 (CUDA) | Source |
|---|---|---|---|
| Hardware acquisition (100 servers) | Significantly lower per-unit cost (projected) † | $30K–$40K per GPU × 8 GPUs/server = $24M–$32M (analyst range) 4 | Analyst consensus; Tenstorrent positioning |
| Cluster power draw | ~800–1,000 kW (100 × 8–10 kW operating) 2 | ~1,200–1,600 kW (100 × 8 GPUs × 1 kW + host) 3 | Tenstorrent spec sheet; NVIDIA datasheet |
| Annual energy cost (at €0.10/kWh, 85% util.) | ~€670K/yr (AGICY estimate) | ~€890K–€1.19M/yr | Calculated from TDP at Cyprus industrial rate |
| Cooling infrastructure (CapEx) | Standard air cooling — included in facility cost | Liquid cooling CDUs, piping, maintenance — $2M–$5M incremental 6 | Industry estimates for 100-server liquid cooling |
| PUE (cooling overhead multiplier) | ~1.3–1.4 (air-cooled) | ~1.1–1.2 (liquid-cooled) | Uptime Institute data centre survey |
| Software / licensing (5-year) | $0 — TT-Metalium (Apache 2.0) | $0–$18M (if AI Enterprise at $4,500/GPU/yr × 800 GPUs × 5 yr) 7 | NVIDIA published pricing; open-source alternative exists |
| Maintenance & ops (5-year, ~10% CapEx/yr) | Lower absolute due to lower CapEx (projected) | $12M–$16M (at 10% of hardware CapEx annually) | Industry standard range |
| Estimated 5-year TCO | Projected 40–60% lower (AGICY estimate) † | $45M–$75M+ (varies significantly by contract) | Model output; NVIDIA range reflects contract variability |
†Tenstorrent has not publicly disclosed per-unit Galaxy pricing. AGICY projections are based on Tenstorrent's stated positioning and our own deployment modelling. These figures will be updated with verified data as AGICY's procurement contracts are finalised.
Beyond Cost: Strategic Considerations
TCO models capture quantifiable costs. But the strategic calculus for AI infrastructure increasingly includes factors that resist neat dollarisation. For European and sovereignty-conscious deployers, these factors can dominate the decision.
Export Control Risk
The US Bureau of Industry and Security (BIS) has progressively tightened restrictions on advanced AI chip exports since October 2022. The January 2025 “AI Diffusion Rule” introduced a three-tier country framework that could affect NVIDIA product availability even for allied-nation customers, depending on end-use classification and data-centre location.8
RISC-V's open instruction-set architecture is not subject to US export licensing in the same way proprietary GPU architectures are. While individual chip implementations may still involve controlled manufacturing processes, the ISA itself — the foundational layer — is open and internationally governed by RISC-V International, a Swiss-domiciled non-profit.
Vendor Concentration
NVIDIA commands an estimated 80–95% of the data-centre AI accelerator market depending on the segment measured.9 This level of vendor concentration creates structural risks: allocation-based pricing, supply shortages during demand spikes, and product roadmap dependencies that customers cannot influence. For organisations where AI inference is a critical capability — not a discretionary experiment — single-vendor dependency is an operational risk, not merely a procurement inconvenience.
Data Sovereignty & Silicon Auditability
RISC-V's open ISA enables a level of silicon transparency impossible with proprietary architectures. For organisations subject to GDPR, financial-services regulations, or government security requirements, the ability to audit inference hardware at the instruction-set level — rather than trusting a vendor's attestation — is a meaningful differentiator.
This is not theoretical. European regulatory frameworks are increasingly demanding supply-chain transparency for critical digital infrastructure. The ability to demonstrate that inference hardware contains no undocumented capabilities or undisclosed data pathways is a compliance advantage.
EU Regulatory Trajectory
Several EU regulatory frameworks are converging to make sovereign infrastructure a compliance consideration, not just a strategic preference:
- Cyber Resilience Act (CRA): Imposes security-by-design requirements on digital products, including hardware. Open-ISA silicon is structurally better positioned for the transparency requirements.
- EU AI Act: High-risk AI systems will face requirements around transparency, traceability, and technical documentation that favour auditable infrastructure stacks.
- NIS2 Directive: Expanded scope for critical infrastructure cybersecurity, with supply-chain security requirements that penalise opaque vendor dependencies.
- European Chips Act: €43 billion investment framework explicitly designed to reduce dependency on non-EU semiconductor supply chains — creating tailwinds for RISC-V adoption in European infrastructure.10
Limitations & Caveats
Intellectual honesty requires acknowledging the significant uncertainties in this analysis. Any TCO model for emerging technology involves assumptions that may not hold.
- Production maturity: Tenstorrent Galaxy is shipping but in early-stage production volumes. Long-term reliability data, sustained performance under production load, and real-world failure rates are not yet available at the scale assumed in this model.
- Software ecosystem:The RISC-V AI software ecosystem is substantially less mature than CUDA. Model optimisation, debugging tools, and framework support — while improving rapidly — do not yet match the depth of NVIDIA's toolchain. Organisations should budget for higher initial software integration effort.
- Projection basis:This analysis is based on AGICY's projected deployment specifications, not production benchmarks from an operating facility. AGICY's data centre is under development; the TCO figures will be validated against actual operational data once the facility is commissioned.
- NVIDIA pricing variability:NVIDIA GPU pricing varies significantly by customer size, geographic region, volume commitment, and competitive dynamics. The analyst-consensus ranges used here may not reflect any specific customer's actual acquisition cost.
- Performance normalisation: This model does not attempt to normalise for tokens-per-second performance differences between the platforms. Independent, at-scale inference benchmarks comparing Galaxy and B200 on identical workloads are not yet publicly available.
- Currency and energy-price assumptions: Energy costs are modelled at €0.10/kWh based on current Cyprus industrial rates. Actual contracted rates and currency fluctuations will affect real-world costs.
Conclusion
RISC-V will not replace NVIDIA overnight. CUDA's ecosystem depth, NVIDIA's training-workload dominance, and the sheer installed base of GPU-optimised software create formidable moats that will persist for years.
But the economics of production AI are not the economics of AI research. Inference workloads — which constitute the overwhelming majority of production AI compute — have different optimisation targets: cost per token, energy efficiency, operational simplicity, and sustained throughput at predictable latency. These are precisely the dimensions where purpose-built RISC-V inference platforms can compete.
For organisations where data sovereignty is a requirement rather than a preference — where vendor independence, supply-chain resilience, and regulatory compliance carry real weight in procurement decisions — the TCO case for RISC-V inference infrastructure is already compelling on a projected basis and will strengthen as the ecosystem matures.
AGICY is building its sovereign AI infrastructure on this thesis: that the next decade of AI deployment will be defined not by who has the fastest training chip, but by who can deliver reliable, auditable, cost-effective inference at scale — under full jurisdictional control.
Sources & References
- 1 NVIDIA data-centre GPU market share estimates from Mercury Research, TechInsights, and JPMorgan equity research (2025–2026). Estimates range from ~80% to 95%+ depending on segment definition.
- 2Tenstorrent Galaxy specifications from Tenstorrent's published product materials and press briefings (2025–2026). 32 Blackhole chips per 6U server, 8–10 kW operating (not nameplate), air-cooled.
- 3NVIDIA B200 specifications from NVIDIA's official Blackwell architecture whitepaper and product datasheets. TDP ~1,000 W per GPU, liquid cooling required.
- 4 NVIDIA GPU pricing ranges from Goldman Sachs, Morgan Stanley, and Barclays semiconductor equity research notes (2025–2026). B200 list pricing estimated at $30,000–$40,000 per GPU; actual contract pricing varies.
- 5 Cyprus industrial electricity rates from the Cyprus Energy Regulatory Authority (CERA) and Eurostat energy price statistics. Long-term contract rates for industrial consumers in the €0.08–€0.12/kWh range.
- 6 Liquid cooling infrastructure costs from Uptime Institute, ASHRAE TC 9.9 guidelines, and Schneider Electric data-centre reference architectures. Per-rack liquid cooling CapEx estimates range widely based on density and cooling approach.
- 7NVIDIA AI Enterprise licensing at $4,500/GPU/year from NVIDIA's published pricing page (as of Q2 2026). Enterprise agreement pricing may differ.
- 8US BIS export control rules: October 2022 (initial China restrictions), October 2023 (expanded scope), January 2025 (“AI Diffusion Rule” with three-tier country framework). Federal Register publications.
- 9NVIDIA market-share estimates from multiple analyst sources. Exact share varies by how “AI accelerator market” is defined and whether cloud-provider custom silicon (Google TPU, AWS Trainium) is included.
- 10 European Chips Act: Regulation (EU) 2023/1781, €43 billion mobilisation target for European semiconductor capacity through 2030.
Methodology Notes & Disclosure
This analysis was prepared by the AGICY Research Team for informational purposes. AGICY.AI is developing a RISC-V-based data-centre facility in Cyprus using Tenstorrent hardware; the company has a commercial interest in the conclusions presented. All AGICY-specific projections are clearly labelled and based on internal deployment modelling, not independently audited production data. NVIDIA cost figures are derived from publicly available analyst reports, published pricing, and vendor documentation — not from AGICY's direct procurement experience. Readers should conduct their own due diligence and consult independent sources before making procurement or investment decisions. This article does not constitute financial or investment advice.
Last updated: July 2026. This is a living document; data will be revised as AGICY's facility enters commissioning and production benchmarks become available.