Bare-Metal GPU Servers for AI Workloads in 2026: Single-Tenant Hardware, ECC RAM, and Pricing Across 10 Configurations from 9 Providers
“Bare metal” has become a marketing term. Three different providers can advertise “bare-metal GPU servers” while one delivers a true single-tenant physical machine, another runs a thin hypervisor layer they call bare metal, and the third splits the same hardware across multiple tenants with vGPU. For AI workloads where performance predictability, GPU memory isolation, and direct driver access matter, only the first definition is genuinely bare metal. The gap between the three categories shows up in measurable ways during sustained training and inference, and the gap shows up in compliance audits, not just performance benchmarks.
Most “best bare-metal GPU” listicles fail at the definitional gate. They mix true single-tenant bare metal with virtualized cloud GPU instances that happen to be sold under a “bare metal” SKU name, then rank them on price-per-GPU-hour as if the underlying hardware experience were equivalent. The result is recommendations that produce different performance characteristics, different cost structures, and different failure modes than buyers expect when they procure based on the article. Worse, the recommendations fail the audit test that regulated buyers (healthcare, finance, defense, EU GDPR scenarios) actually need to pass.
This guide applies a strict definitional test: a Bare-Metal GPU Server provider is included only if it offers true single-tenant physical hardware with no hypervisor layer between the customer and the GPU. By that standard, 10 configurations from 9 distinct providers qualify across four infrastructure tiers, from EUR 184/month single-GPU entry servers to enterprise H100 clusters with 3,200 Gbps InfiniBand fabric. Hostline is one of the providers compared, with methodology and disclosure documented at the end.
The framework sections cover what “true bare metal” actually means in technical terms, what virtualization overhead measurably costs AI workloads, when bare metal beats cloud for AI specifically, and how to match workload type (training, inference, fine-tuning, RAG) to the appropriate bare-metal tier.
Across the providers evaluated, the EU value tier offers the most accessible entry points for AI workloads. Hetzner GEX44 at EUR 184/month delivers single-GPU bare metal with 20 GB GDDR6 ECC and native FP8 from Falkenstein, Germany [4]. Hostline provides multi-GPU bare-metal configurations (single A4000, dual A5000, triple A5000) at fixed monthly EUR pricing from EUR 360/month with ECC system RAM at every tier and zero egress fees from Vilnius, Lithuania [3]. Hetzner GEX131 at EUR 889/month per Hetzner’s monthly commitment pricing (or EUR 1.42 per hour on the separate on-demand tier, which works out to approximately EUR 1,037 per month at 24/7 utilization, reflecting an on-demand premium over the monthly commitment) provides 96 GB GDDR7 Blackwell with FP4 support and NVMe Gen4 on a single GPU [4]. Cherry Servers offers broader GPU selection (A100, A40, A10, A2) with hourly billing flexibility and 100 TB/month free egress from Lithuania [6][7].
In the global and US tier, Latitude.sh delivers API-first bare metal with H100 PCIe, RTX PRO 6000, and L40S across approximately 25 global locations with hourly and monthly billing; Latitude.sh was acquired by Megaport Limited (ASX: MP1) on November 27, 2025 per PitchBook and the Megaport press release, integrating Latitude.sh’s compute platform with Megaport’s global private connectivity fabric [15][16]. Liquid Web offers H100 NVL bare metal from US data centers with explicit “up to 15% better than virtualized” positioning [18] and 10 TB included egress [17]. Oracle OCI is the only hyperscaler among the providers evaluated with true bare metal, scaling from single 8-GPU nodes to 131,072-GPU Superclusters spanning H100, H200, B200, and MI300X [19].
In the specialized AI tier, Voltage Park (now operating under the Lightning AI name following the January 2026 merger completion [24][25]) provides single-tenant HGX H100 bare-metal clusters with 3,200 Gbps Quantum-2 InfiniBand starting at $1.99/hr on-demand [23]. FluidStack builds single-tenant clusters from 8 to 30,000 GPUs (selected by Anthropic as the build partner for its announced $50 billion US data center investment, per Anthropic’s November 2025 announcement [30] and the December 2025 Hut 8 site partnership [33]) with H100, H200, B200, and GB200 options [29]. TensorWave is the only AMD-only bare-metal provider here, with MI355X starting at $2.85/GPU-hr, MI325X at $1.95/GPU-hr, and MI300X options [34].

How This Comparison Was Built
Providers were evaluated against a strict definitional test: a provider is included only if it offers true single-tenant physical GPU hardware with no hypervisor layer between the customer and the GPU. This test excludes virtualized cloud GPU offerings (Lambda Cloud GPU instances, RunPod, AWS P5, Azure ND-series, Google A3/A4, Nebius cloud instances), GPU marketplaces (Vast.ai), and providers that market “bare metal” but run lightweight hypervisor virtualization (Crusoe, per SemiAnalysis ClusterMAX evaluations). Providers offering both virtualized and true bare-metal SKUs are included only for their true bare-metal lines.
Evaluation dimensions are GPU SKUs and VRAM per card, CPU and RAM specifications, storage type and capacity, networking bandwidth, data center locations and certifications (ISO 27001, SOC 2, HIPAA, GDPR, Uptime Institute Tier), system RAM ECC availability, multi-GPU-in-chassis capability, IPMI/iDRAC/BMC remote management, egress policy, and monthly or hourly pricing in May 2026. Where a vendor publishes a performance claim (Liquid Web’s “up to 15% better than virtualized,” TensorWave benchmarks, OVHcloud’s “64% cost reduction”), the commercial interest is flagged explicitly in the provider section.
Providers are grouped into four infrastructure tiers: EU value bare-metal, global and US bare-metal, hyperscaler bare-metal, and specialized AI bare-metal. Within each tier, providers are presented in a narrative ordering that walks from entry-level capability to higher-spec configurations; the precise per-attribute comparison appears in the comparison table rather than the section ordering. No provider is ranked #1 overall. Hostline is the publisher of this article and appears at position 2 within the EU value tier. Where Hostline’s hardware falls short of competitors on a specific dimension, this is stated directly. The table contains 10 rows. Hetzner is evaluated across two distinct SKU tiers (GEX44 and GEX131), each with materially different VRAM and pricing characteristics worth comparing individually.
What “True Bare Metal” Actually Means
True bare metal in 2026 is defined by what the customer can do, not by what the marketing page says. A genuinely bare-metal GPU server gives the customer direct physical access to the hardware: no hypervisor layer between the operating system and the GPU, no shared tenancy on the physical machine, BIOS and firmware access for memory speed and NUMA tuning, direct PCIe access to the GPU without virtualization shims, custom kernel driver installation, and unrestricted access to nvidia-smi or rocm-smi for hardware diagnostics.
Three categories of “bare metal” exist in current vendor marketing. The first is true bare metal: the customer rents an entire physical server with exclusive hardware access. Hostline, Hetzner GEX, Cherry Servers, Oracle OCI Bare Metal shapes, Voltage Park dedicated reserve (now Lightning AI dedicated reserve), FluidStack single-tenant clusters, TensorWave bare metal, Liquid Web Metal, and Latitude.sh Metal SKUs all qualify. The second is “bare metal as a service” with a lightweight hypervisor layer, where some Vultr offerings, parts of Crusoe’s catalog, and certain neocloud “bare metal” SKUs fall. SemiAnalysis ClusterMAX evaluations have specifically flagged Crusoe as running cloud-hypervisor VMs rather than true bare metal [40]. The third category is virtualized cloud GPU sold under premium “bare metal” branding, excluded from this comparison entirely.
The hardware access difference matters concretely. With true bare metal, a customer can install a custom CUDA driver version not yet packaged for managed clouds, configure GPUDirect RDMA between cards without virtualization breaking the topology, run BIOS-level memory tuning that affects training throughput on memory-bandwidth-bound workloads, and inspect physical PCIe topology with lspci -tv to verify NUMA pinning before launching multi-GPU jobs. With a hypervisor in the way, some operations are restricted, some return virtualized values rather than physical ones, and some work but produce different results than the bare-metal equivalent.
For AI workloads specifically, GPU memory isolation prevents data leakage in multi-tenant scenarios, which becomes a compliance requirement for healthcare data under HIPAA, financial data under PCI DSS, and certain EU GDPR processing scenarios. Predictable PCIe topology matters for tensor parallelism across multi-GPU configurations. Hardware-specific tuning (NVLink bridge configuration, NUMA node pinning, BIOS-level memory frequency) requires bare metal because these settings either do not exist in virtualized environments or are managed centrally by the cloud provider.
What Virtualization Overhead Actually Costs AI Workloads
The literature on virtualization overhead for GPU workloads spans 2 percent to 25 percent depending on workload type, GPU configuration, and tuning quality. Honest synthesis requires acknowledging this range rather than citing a single number from a vendor marketing page.
At the low end, Principled Technologies’ controlled benchmarks of virtualized NVIDIA A100 against bare-metal A100 found the virtualized configuration reached approximately 97.5 percent of bare-metal performance on standard ML training benchmarks, representing roughly 2.5 percent overhead [36]. Academic HPC and deep learning studies typically report under 10 percent overhead for full-card vGPU configurations on training workloads where PCIe passthrough is correctly configured. A Springer benchmark study measured approximately 5 percent container overhead, approximately 10 percent VM overhead, and approximately 15 percent combined when both layers are present [37].
At the high end, vendor marketing materials cite 15 to 25 percent overhead, with Hivelocity referring to a “virtualization tax” of 5 to 25 percent [38] and Aethir citing real-world deployments at 15 to 25 percent [39]. Liquid Web’s own marketing claims “up to 15% better GPU performance over virtualized environments” for its Metal line [18]. These figures are commercial in nature and should be treated as upper bounds for poorly tuned environments rather than typical results.
Where the overhead is largest matters more than the average. Memory-bandwidth-bound inference workloads (LLM token generation, where each token requires reading model weights and KV cache from VRAM into the tensor cores) suffer disproportionately because virtualization adds latency to every memory transaction. Multi-tenant time-sliced vGPU configurations can show overhead well above 25 percent during contention. By contrast, single-GPU training with PCIe passthrough on compute-bound workloads with large batch sizes shows the smallest overhead, often below 5 percent.
The practical buyer implication is that virtualization overhead matters most when utilization is already high enough that bare metal wins on cost. For a 24/7 inference workload at 80 percent sustained utilization on an 8B model, a 10 percent throughput penalty represents roughly $1,000 to $3,000 per year in additional cloud costs versus equivalent bare metal, on top of the cloud-vs-bare-metal cost premium itself. The overhead is rarely the decisive factor in choosing bare metal over cloud, but it compounds with the cost premium and shifts the break-even utilization threshold lower than the cost analysis alone would suggest.
When Bare Metal Beats Cloud for AI Workloads
The economic break-even between bare metal and cloud is approximately 60 to 70 percent sustained utilization on H100-class hardware, with the threshold shifting lower for lower-cost Ampere bare metal. Cherry Servers’ own published guidance on bare-metal economics centers on a similar threshold per its public bare-metal economics blog, with break-even typically reached within 6 to 12 months [9]. Below approximately 40 percent utilization, on-demand cloud GPU pricing wins because idle time costs nothing.
Four AI workload patterns map to four positions on the utilization curve. Sustained inference serving (production chatbots, API endpoints, batch inference pipelines) typically runs at 60 to 95 percent utilization, the ideal bare-metal fit. Continuous fine-tuning pipelines (daily or weekly LoRA runs on customer data) run at 70 to 90 percent during the runs themselves but may be idle between runs, putting them in the moderate-fit zone. One-off training runs (model development, experimentation) typically achieve 20 to 50 percent utilization across calendar time, where cloud burst capacity wins. RAG production services (vector database queries plus embeddings plus LLM inference) typically run at 40 to 80 percent utilization, where bare metal often wins but the calculation depends on traffic patterns.
A worked example using May 2026 prices makes the break-even concrete. Hostline’s dual RTX A5000 configuration costs EUR 903 per month, approximately $985 at May 2026 exchange rates of 1.09 USD/EUR. A price-comparable cloud option is two A100 40 GB GPUs on Lambda Cloud at $1.29 per GPU-hour, accumulating approximately $1,884 per month at 24/7 utilization. The cost comparison is meaningful, but the cards are not performance-equivalent: the data-center A100 carries HBM2 memory at 1.55 TB/s bandwidth and substantially higher tensor throughput than the workstation-class Ampere A5000 at 768 GB/s GDDR6. The bare-metal break-even occurs at approximately 52 percent utilization (Hostline’s $985 monthly cost divided by Lambda’s $1,884 monthly cost at 24/7), lower than the 60 to 70 percent threshold for higher-end GPUs because the A-series Hostline hardware costs less in absolute terms. The choice between Hostline dual A5000 and cloud 2x A100 depends on whether the workload requires A100-class memory bandwidth (vision-language models, FP8 acceleration) or fits comfortably on A5000-class hardware (most production YOLO, ResNet, EfficientNet, classical Mask R-CNN, OCR).
The compliance and isolation factor shifts the calculation independently of pure economics. Regulated AI workloads (healthcare under HIPAA, financial data under PCI DSS, defense and ITAR-controlled work, certain EU GDPR processing scenarios) often require demonstrable hardware-level isolation that virtualized cloud cannot provide regardless of compliance attestations. For these workloads, bare metal is the legally compliant option, and the cost comparison does not apply because cloud is not viable. Hidden cloud costs also shift the calculation: hyperscaler egress fees add 5 to 15 percent to effective cost, and bare-metal providers with zero egress (Hostline, Voltage Park / Lightning AI, Cherry Servers’ 100 TB included, Latitude.sh’s 20 TB included per server) eliminate this entirely.
Matching AI Workloads to Bare-Metal Tiers
Different AI workloads have different bare-metal requirements, and the right provider depends on workload pattern as much as on GPU generation.
LoRA and QLoRA fine-tuning of 7B to 13B parameter models requires 16 to 24 GB of VRAM. Hetzner GEX44 (EUR 184/month, 20 GB), Hostline single A4000 (EUR 360/month, 16 GB), and Cherry Servers A10 ($0.479/hr, 24 GB, subject to waitlist availability) cover this tier at the lowest cost. Full fine-tuning of 30B to 70B models with NVLink requires 80 GB or more per GPU and multi-GPU NVLink, narrowing the options to Oracle OCI BM.GPU.H100.8, FluidStack single-tenant H100/H200 clusters, and Voltage Park / Lightning AI dedicated reserve HGX H100. Hostline cannot serve this tier because its Ampere-generation RTX A-series hardware lacks data center H100 SXM with NVLink.
Sustained inference serving on 7B to 32B quantized models works in the 16 to 48 GB VRAM range, where memory bandwidth matters more than raw compute. Hostline provides fixed monthly EUR pricing, Cherry Servers A40 or A10 add hourly billing flexibility, and Liquid Web L40S covers US data residency. High-VRAM inference on 70B+ models at INT4 with long context windows requires 96 to 192 GB per GPU: Hetzner GEX131 (96 GB at EUR 889/month), Oracle OCI B200 or MI300X 8-GPU shapes, TensorWave MI355X (288 GB) or MI325X (256 GB), and Liquid Web H100 NVL (94 GB).
Pretraining and frontier training of 100B+ models across multiple nodes requires InfiniBand fabric and 8 to 64+ GPUs. Voltage Park / Lightning AI scales to 4,064 GPUs per dedicated cluster across a combined fleet of 35,000+ GPUs post-merger, FluidStack covers 8 to 30,000 GPUs, and Oracle OCI Superclusters scale to 131,072 GPUs. RAG production services place demand on storage I/O and host RAM as much as on the GPU, where Hostline’s triple A5000 (256 GB DDR4 ECC, 2x 1.92 TB SSDs), Hetzner GEX131 (NVMe Gen4), and Cherry Servers EPYC base (up to 1024 GB ECC RAM, 80 TB storage) all fit.
Best Bare Metal GPU Server Providers
| Provider | Tier | Top GPU | VRAM/GPU | Multi-GPU Chassis | System RAM ECC | Network | Locations | Pricing | Egress |
|---|---|---|---|---|---|---|---|---|---|
| Hetzner GEX44 | EU value bare-metal | RTX 4000 SFF Ada | Y20 GB GDDR6 (VRAM ECC; system RAM non-ECC) | No (single GPU) | No (i5-13500, consumer) | 1 Gbps | DE | EUR 184/mo | Unmetered |
| Hostline | EU value bare-metal | RTX A5000 | 24 GB GDDR6 ECC | Yes (up to 3) | Yes | 1 Gbps | LT | EUR 360-1,220/mo | Unmetered |
| Hetzner GEX131 | EU value bare-metal | RTX PRO 6000 Blackwell | 96 GB GDDR7 ECC | No (single GPU) | DCV, RDP, Anyware | 1 Gbps (10 Gbps opt) | DE | EUR 889/mo (monthly commit); EUR 1.42/hr (on-demand) | Unmetered (10G uplink: over 20 TB outbound EUR 1/TB) |
| Cherry Servers | EU value bare-metal | A100 80GB | 80 GB HBM2e | Yes (EPYC base up to 2) | Yes | Up to 10 Gbps | LT, NL, DE, SE, US, SG, JP | $0.22-2.18/hr | 100 TB free |
| Latitude.sh | Global API-first bare-metal | H100 PCIe | 80 GB HBM2e | Yes | Yes | Up to 100 Gbps | ~25 locations | from ~$1.70/hr (H100 PCIe) | 20 TB free |
| Liquid Web | US bare-metal | H200 NVL | 141 GB HBM3e | Yes (up to 2x H200 NVL) | Yes | 10 Gbps | US | $1.07/hr (L4) to $4.38/hr (H200 NVL single); up to $7.19/hr (dual H200 NVL) | 10 TB free |
| Oracle OCI | Hyperscaler bare-metal | H100/H200/B200/ B300/MI300X | 80-288 GB (per NVIDIA spec) | Yes (8-GPU shapes) | Yes | RoCE v2 RDMA | Global | ~$10/GPU-hr | Free (global, since Feb 2026) |
| Voltage Park / Lightning AI | Specialized AI bare-metal | HGX H100 SXM | 80 GB HBM3 | Yes (HGX nodes) | Yes | 3,200 Gbps IB | US (TX/VA/WA/UT) | From $1.99/hr | Unmetered |
| FluidStack | Specialized AI bare-metal | L4H100/H200/B200/GB200 | 80-192 GB | Yes (clusters) | Yes | InfiniBand | Global | Quote-based | Unmetered |
| TensorWave | Specialized AI bare-metal (AMD) | MI355X/MI325X/MI300X | 192-288 GB HBM3/HBM3E | Yes | Yes | High-speed | US (Las Vegas) | $1.95-2.85/GPU-hr | Unmetered |
EU Value Bare-Metal: Hetzner GEX44

source: hetzner.com
Hetzner Online operates its GPU servers from its German data centers in Falkenstein and Nuremberg, ISO 27001 certified, GDPR compliant, running on 100 percent green energy [4]. Among the GPU models, the GEX44 is offered only in the Falkenstein (FSN1) data center, while the GEX131 is available in both Falkenstein and Nuremberg [4]. The GEX44 is the entry tier and the cheapest serious bare-metal GPU server among the providers evaluated.
The GEX44 ships an NVIDIA RTX 4000 SFF Ada Generation GPU with 20 GB GDDR6 ECC VRAM and native FP8 tensor core support, paired with an Intel Core i5-13500, 64 GB DDR4 (non-ECC, since the i5-13500 is a consumer Raptor Lake part without system ECC support), and 2x 1.92 TB NVMe Gen3 SSDs in software RAID 1. Pricing is EUR 184 per month with a EUR 79 setup fee, confirmed by multiple May 2026 third-party sources [4]. The 1 Gbps networking is unmetered, with no traffic caps. IPMI remote management is included.
For an AI buyer, the GEX44 covers LoRA fine-tuning of 7B to 13B models, sustained inference on 7B to 14B INT4 quantized models, and RAG workloads where 20 GB of VRAM is sufficient. The native FP8 support on the Ada Lovelace architecture is the key differentiator versus older Ampere hardware, providing throughput acceleration on quantized inference and FP8 mixed-precision training. The constraint is the single-GPU configuration: Hetzner explicitly states that GEX servers “each have one GPU and cannot be configured with multiple GPUs,” ruling out multi-GPU parallelism within a single chassis [4].
Strengths: lowest monthly cost for true bare metal among providers evaluated at EUR 184/month; native FP8 tensor core support on Ada Lovelace; ECC GDDR6 VRAM on the GPU; NVMe Gen3 storage in software RAID 1; ISO 27001 and GDPR certified; 100% green energy; unmetered 1 Gbps networking.
Limitations: single GPU only with no multi-GPU configurations in a single chassis; 20 GB VRAM caps the workload at roughly 14B parameters at FP16 or 32B at INT4; non-ECC system DDR4 RAM (the i5-13500 is consumer Raptor Lake without system ECC support, though GPU VRAM is ECC); EUR 79 setup fee on monthly subscriptions; setup takes 1 to 3 business days.
EU Value Bare-Metal: Hostline

source: hostline.io
Per its published documentation, Hostline operates dedicated bare-metal GPU servers from a data center built to Tier III standards in Vilnius, Lithuania, providing fixed monthly EUR pricing, ECC system RAM at every tier, and no egress fees [3].The infrastructure operates with redundant N+1 network connectivity and power delivery systems and full GDPR compliance for EU data residency on personal data per Hostline’s published infrastructure documentation. Hostline is owned and operated by HOSTLINE UAB (Lithuanian company registration number 302660481).
Three SKUs are available, all on the Intel Xeon Gold 6130 platform. The entry plan pairs a single NVIDIA RTX A4000 (16 GB GDDR6 ECC) with 64 GB DDR4 ECC RAM and 2x 960 GB SSDs at EUR 360 per month. The mid-tier plan provides two NVIDIA RTX A5000 GPUs (48 GB GDDR6 ECC aggregate) with two Xeon Gold 6130 CPUs, 128 GB DDR4 ECC RAM, and 2x 1.92 TB SSDs at EUR 903 per month. The top configuration adds a third RTX A5000 (72 GB GDDR6 ECC aggregate) and scales system RAM to 256 GB DDR4 ECC at EUR 1,220 per month [3]. All plans include iDRAC 9 Enterprise for out-of-band remote management, full root access, RAID 0/1/10 storage options, DDoS protection, and zero egress fees. Dedicated servers provision within one business day.
Hostline’s position in the EU value tier is defined by a specific combination that does not appear in any other provider evaluated here: multi-GPU bare-metal configurations (dual and triple RTX A5000) at fixed monthly EUR pricing with no on-demand tier, ECC system RAM at the entry-level GPU plan, and zero egress fees. Each of these properties exists at other providers individually (Hetzner publishes fixed monthly EUR pricing on GEX44 and GEX131; Cherry Servers’ EPYC base supports up to two GPUs in a chassis; Liquid Web, Latitude.sh, and Oracle all use server-class CPUs with ECC system RAM on their respective tiers; multiple providers offer free or generous-allowance egress). The combination at the EU value tier price point is what is distinct: Hetzner’s GEX line is strictly single-GPU per Hetzner’s own GEX product page; Cherry Servers’ multi-GPU EPYC base is priced in USD, with several GPU SKUs on pre-order or waitlist, rather than at fixed monthly EUR pricing; Hetzner GEX44’s entry tier uses non-ECC consumer DDR4 system RAM (the i5-13500 has no system ECC support) while Hostline’s entry-level RTX A4000 plan ships DDR4 ECC.
For inference, the 16 GB RTX A4000 handles Llama 3.1 8B, Mistral 7B, Qwen 3 8B, and Gemma 2 9B at FP16, plus Phi-4 14B at INT4. The 24 GB RTX A5000 extends coverage to quantized 27B to 32B models with KV cache headroom for 8B models at long contexts. The dual A5000 enables two independent inference replicas behind a load balancer, with each replica serving its own model on one card. The triple A5000 enables three independent replicas. Multi-GPU tensor-parallel serving of a single 70B model is not a viable configuration on three RTX A5000 cards: vLLM’s tensor_parallel_size must evenly divide both the model’s query-head count and KV-head count, and Llama-class 70B models with 64 query heads and 8 KV heads cannot be split across three GPUs [1]. Additionally, the NVIDIA RTX A5000 NVLink bridge is documented as a 2-way bridge per NVIDIA’s product page [2], so three A5000 cards cannot be NVLinked together. For 70B INT4 serving across two of the three cards, TP=2 with the third card serving a separate workload is a valid configuration, provided the customer confirms NVLink bridge availability with Hostline before deployment. Without NVLink, tensor-parallel inference over PCIe Gen3 (32 GB/s bidirectional) carries a meaningful throughput penalty.
At EUR 903 per month for the dual A5000 configuration (approximately $985 at May 2026 exchange rates of 1.09 USD/EUR), the effective rate works out to approximately EUR 0.62 per GPU-hour (approximately $0.67) at 24/7 operation. A price-comparable cloud option is two A100 40 GB GPUs on Lambda Cloud at $1.29 per GPU-hour, accumulating approximately $1,884 per month at 24/7 utilization. The cost comparison is meaningful, but the cards are not performance-equivalent: the data-center A100 carries HBM2 memory at 1.55 TB/s bandwidth and substantially higher tensor throughput than the workstation-class Ampere A5000 at 768 GB/s GDDR6. For sustained 24/7 CV inference, fine-tuning of 7B-13B models, and RAG workloads from the EU under GDPR data residency requirements, the dual or triple A5000 configurations cover the workload at the lowest sustained monthly cost among true bare-metal providers evaluated here. For workloads that require A100-class memory bandwidth (large vision-language models, FP8 acceleration, compute-heavy fine-tuning), cloud A100 or Hopper-class hardware is the right hardware choice regardless of the cost differential.
Strengths: lowest sustained monthly cost for multi-GPU bare metal among providers evaluated at EUR 903/month for dual A5000 (48 GB aggregate VRAM); fixed monthly EUR pricing with no on-demand premium tier (Hetzner GEX131 publishes both a monthly commitment and a separate on-demand hourly tier at a roughly 17 percent premium; Hostline publishes only the monthly tier); the only multi-GPU-in-chassis bare-metal configuration in the EU value tier at fixed monthly EUR pricing across the providers evaluated; ECC DDR4 system RAM at the EUR 360/month entry-level plan (a distinction from Hetzner GEX44, whose i5-13500 consumer Raptor Lake CPU does not support system ECC and so ships non-ECC DDR4 at the same price band); EU/GDPR data residency in Vilnius, Lithuania; data center built to Tier III standards per Hostline’s published documentation; full root access with iDRAC 9 Enterprise on all plans; zero egress fees on all GPU plans; 256 GB DDR4 ECC system RAM on the triple A5000 plan; DDoS protection included per Hostline’s product page; redundant N+1 network connectivity and power delivery per Hostline’s published infrastructure documentation.
Limitations: Ampere-generation RTX A4000 and A5000 GPUs without native FP8 tensor core support, yielding lower throughput per watt than Ada Lovelace, Hopper, or Blackwell GPUs on FP8-optimized workloads; Intel Xeon Gold 6130 CPUs (Skylake-SP, 2017) with PCIe Gen3 bandwidth constraints; SSDs are listed on the Hostline product page without explicit NVMe Gen4 designation, suggesting SATA or earlier NVMe generations relative to Hetzner GEX131’s published NVMe Gen4 (verify drive interface directly with Hostline before specifying for I/O-bound large training); 1 Gbps Ethernet networking limits concurrent API throughput and rules out multi-node distributed training; 24 GB maximum VRAM per GPU rules out full fine-tuning above approximately 32B parameters without aggressive quantization; no publicly listed SOC 2 or ISO 27001 attestation on Hostline’s site beyond the Tier III standards its data center is built to, per Hostline’s own documentation; NVLink bridge documented as 2-way only on RTX A5000 per NVIDIA, ruling out three-card NVLinked tensor parallelism; single location (Vilnius) with no global presence.
EU Value Bare-Metal: Hetzner GEX131

source: hetzner.com
The GEX131 is Hetzner’s flagship single-GPU bare-metal server, shipping an NVIDIA RTX PRO 6000 Blackwell Max-Q with 96 GB GDDR7 ECC at EUR 889 per month on the monthly commitment tier [4]. Hetzner separately publishes an on-demand hourly rate of EUR 1.42 per hour, which works out to approximately EUR 1,037 per month at 24/7 utilization, reflecting an approximately 17 percent premium over the monthly commitment in exchange for hourly flexibility. The system platform is an Intel Xeon Gold 5412U (24 cores, Sapphire Rapids), 256 GB DDR5 ECC expandable to 768 GB, and 2x 960 GB NVMe Gen4 SSDs. Nuremberg and Falkenstein locations, ISO 27001 certified, GDPR compliant, 100 percent green energy. Hetzner lists the GEX131 in multiple RAM and storage configurations from the EUR 889 per month base [4]. An optional 10 Gbps uplink is available, and when it is added Hetzner’s otherwise unlimited traffic policy caps outbound traffic at 20 TB per month with usage beyond that billed at EUR 1.00 per TB; the default 1 Gbps uplink stays fully unmetered [4].
The 96 GB VRAM is the largest single-GPU capacity in the EU value tier here. Llama 3.3 70B at INT4 fits with substantial KV cache headroom, Llama 3.3 70B at INT8 fits with moderate headroom, and 30B to 40B models at FP16 fit comfortably. The Blackwell generation adds native FP4 tensor core support in addition to FP8, enabling further quantization gains on supported frameworks. Puget Systems testing has shown the Max-Q variant runs approximately 5 to 14 percent slower than the 600W Workstation Edition of the same chip while drawing 300W rather than 600W, a favorable trade-off for sustained inference loads where power efficiency matters [5].
Strengths: 96 GB GDDR7 ECC VRAM in a single bare-metal GPU, the highest single-GPU VRAM in the EU value tier; Blackwell-generation FP4 and FP8 tensor core support; NVMe Gen4 storage in RAID configuration; DDR5 ECC expandable to 768 GB; ISO 27001 and GDPR certified; optional 10 Gbps uplink; monthly commitment at EUR 889/mo or on-demand hourly billing at EUR 1.42/hr.
Limitations: single GPU only; 1 Gbps default networking (10 Gbps optional); pricing requires EUR 79 setup fee on monthly subscriptions; no published SLA percentage; pricing discrepancies reported between Hetzner press releases and live forum data (verify live storefront price before purchasing).
EU/Global Value Bare-Metal: Cherry Servers

source: cherryservers.com
Cherry Servers operates bare-metal infrastructure from Lithuania (Šiauliai), with partner data centers in the Netherlands (Amsterdam), Germany (Frankfurt), Sweden (Stockholm), the US (Chicago), Singapore, and Tokyo (the Tokyo location went live in May 2026 per Cherry Servers’ published announcement [8]). GPU servers are restricted to the Lithuania facility, which is certified ISO 27001, ISO 22301, ISO 9001, SOC 1 Type II, and SOC 2 Type II per Cherry Servers’ published location page [7]. The company offers true bare metal with no setup fees and hourly, monthly, or annual billing.
The GPU lineup is broader than Hostline’s or Hetzner’s: NVIDIA A100 80GB starting at $2.18/hr or $1,530.17/month (pre-order), A40 48GB starting at $0.74/hr or $436.30/month (pre-order), A16 64GB starting at $0.498/hr or $290.52/month (waitlist), A10 24GB starting at $0.479/hr or $279.72/month (waitlist), A2 16GB starting at $0.22/hr or $128.52/month, and Tesla P4 at $0.157/hr [6]. GPUs are added to a Custom Dedicated Server base, with base servers starting at $158.13/month or $0.318/hr. CPU options include AMD Ryzen 7700X, AMD EPYC 7402P, and Intel Xeon Gold 6230R. The EPYC base configuration supports up to two GPUs per chassis; Ryzen and Intel bases support one. Up to 1024 GB ECC DDR4 RAM and up to 80 TB storage are available. Networking includes up to 10 Gbps uplinks with 100 TB/month free egress, overage at approximately EUR 0.5/TB.
Cherry Servers’ hourly billing and broader GPU menu (A100 down to A2) make it more flexible than Hostline for buyers who need lower entry pricing on a smaller GPU (A2 at $128.52/month) or higher VRAM than Hostline offers (A100 at 80 GB versus Hostline’s 24 GB ceiling). The constraint is availability: several SKUs are pre-order or waitlist as of May 2026, and high-end GPUs (A100, A40) can have multi-week lead times.
Strengths: broadest GPU menu in the EU value tier (A100, A40, A16, A10, A2, P4); hourly, monthly, or annual billing flexibility; no setup fees; 100 TB/month free egress; EPYC base supports up to 2 GPUs in a chassis; up to 1024 GB ECC DDR4 and 80 TB storage on EPYC base; 7 global data center locations company-wide (Tokyo added May 2026).
Limitations: GPU servers restricted to Lithuania; several GPU SKUs (A100, A40, A16, A10) are pre-order or waitlist with multi-week lead times; no NVLink documented on multi-GPU configurations; 24-72 hour GPU server deployment; no Hopper, Blackwell, or MI300X options.
Global API-First Bare-Metal: Latitude.sh

source: latitude.sh
Latitude.sh provides API-first bare-metal infrastructure across approximately 25 global locations spanning North America, Europe, Latin America, and Asia-Pacific (Sydney was added in 2025 per the company’s announcements) [10]. In November 2025, Latitude.sh was acquired by Megaport Limited (ASX: MP1), the Brisbane-based Network-as-a-Service provider, integrating Latitude.sh’s compute platform with Megaport’s global private connectivity fabric per the Megaport press release [15][16]. The company centers its offering on dedicated bare-metal servers with API-driven provisioning, hourly and monthly billing, and a documented REST API for server lifecycle management.
GPU SKUs include the NVIDIA H100 PCIe, RTX 6000 Ada, RTX PRO 6000, and L40S, with dual 100 Gbps networking on multi-GPU nodes. Pricing varies by configuration and region, with hourly and monthly billing published on Latitude.sh’s own pricing page [12]; independent trackers in 2026 placed Latitude’s single H100 PCIe in the range of roughly $1.70 to $3.40 per hour depending on region and availability. Atlantic.net reports a single A100 80GB server at approximately $2,234.88/month and a single H100 PCIe server at approximately $2,995.20/month on fixed monthly billing as of early 2026 [13]. Each server includes 20 TB free outbound bandwidth per month, with additional bandwidth in 10 TB packages priced by region ($1.25/TB in US/DE/UK/NL, $5.40/TB in LATAM/APAC). Bare-metal provisioning takes 3 to 10 minutes per the vendor’s own documentation [14], faster than typical bare-metal provisioning at other providers but slower than virtual instance startup times of 30 to 60 seconds. Service level agreements are documented at up to 99.9 percent per Latitude.sh’s published blog [11].
The API-first model and 20-location global footprint make Latitude.sh useful for buyers who need geographic distribution combined with true bare metal. The constraints are GPU breadth and supply: Latitude offers H100 PCIe but not H100 SXM, no InfiniBand-connected multi-GPU configurations for distributed training, and frequent H100 stockouts. Billing rounds to the full hour rather than per-second or per-minute, which adds cost on short training jobs. The platform suits enterprise ML teams running production training where reproducibility and performance consistency matter, regulated industries needing guaranteed hardware isolation, and teams with GDPR or data-residency requirements that need EU bare metal at scale.
Strengths: API-first provisioning with documented REST API; approximately 25 global locations across North America, Europe, Latin America, and Asia-Pacific; dual 100 Gbps networking on multi-GPU nodes; bare-metal provisioning in 3 to 10 minutes; 20 TB free outbound bandwidth per server per month; hourly and monthly billing options; up to 99.9 percent uptime SLA per Latitude.sh’s published blog; integration with Megaport’s global private connectivity fabric post-acquisition.
Limitations: H100 PCIe only with no H100 SXM available; no InfiniBand-connected multi-GPU configurations for distributed training; H100 supply frequently constrained with reported stockouts; billing rounds to full hour rather than per-second or per-minute; pricing at the H100 tier is above specialized AI clouds (Voltage Park / Lightning AI H100 at $1.99/hr, below Latitude’s H100 PCIe rates); H100 limited to a handful of regions (Dallas, São Paulo, Tokyo per Spheron analysis [14]).
US Bare-Metal: Liquid Web

source: liquidweb.com
Liquid Web operates bare-metal GPU infrastructure from US data centers under its “Metal” line, with both portal-based and API-driven provisioning. The company markets its GPU servers with explicit positioning against virtualized alternatives, claiming “up to 15% better GPU performance over virtualized environments” (vendor-published, not independently verified).
The GPU lineup spans the NVIDIA L4 Ada (24 GB) at $1.07/hr, NVIDIA L40S (48 GB) at $1.92/hr, NVIDIA H100 NVL (94 GB) at $3.97/hr, and a dual H100 NVL configuration at $6.85/hr [17]. The L4 entry tier pairs the GPU with dual AMD EPYC 9124 CPUs, 128 GB DDR5 RAM, and a 1.92 TB NVMe SSD. The H100 NVL configurations target high-VRAM inference workloads (94 GB per card) and full fine-tuning at smaller scale. Each server includes 10 TB of outbound bandwidth.
The Liquid Web positioning is North America-centric: GPU bare-metal hosting is documented from US data centers (Lansing, Michigan and Phoenix, Arizona), with the company’s broader infrastructure also including Amsterdam, London, San Jose, Ashburn, and Sydney across the dedicated server portfolio. The GPU Metal line is not explicitly documented as available from EU data centers as of May 2026; verify with Liquid Web sales for current EU GPU availability. For US-based AI workloads where 94 GB single-GPU inference, dual H100 NVL configurations, or L40S production serving are required, Liquid Web competes directly with Latitude.sh and provides a polished managed-bare-metal experience with portal and API options. The vendor’s “up to 15% better than virtualized” claim falls within the upper bound of the virtualization overhead literature (2 to 25 percent range), so the figure is plausible for poorly tuned virtualized environments but may overstate the benefit for well-configured passthrough VMs.
Strengths: H100 NVL 94 GB single-GPU and dual H100 NVL configurations for high-VRAM inference; dual AMD EPYC 9124 CPUs on the L4 tier; DDR5 RAM and NVMe storage standard; bare-metal API and portal provisioning; 10 TB included egress per server.
Limitations: GPU bare-metal availability documented from US data centers only with no published EU or APAC GPU presence on the Metal line; “up to 15% better than virtualized” performance claim is vendor-published and not independently verified; no InfiniBand-connected multi-node configurations; no NVLink-equipped HGX H100 SXM options (only H100 NVL); pricing at the H100 tier is above specialized AI clouds (Voltage Park / Lightning AI H100 at $1.99/hr versus Liquid Web H100 NVL at $3.97/hr).
Hyperscaler Bare-Metal: Oracle OCI

source: oracle.com
Oracle Cloud Infrastructure is the only hyperscaler among providers evaluated offering true bare-metal GPU instances with no hypervisor layer. The OCI bare-metal shapes provide direct hardware access for AI workloads requiring zero virtualization overhead, hyperscaler-grade compliance (SOC 2, ISO 27001, HIPAA, PCI), and the scale to provision Superclusters up to 131,072 GPUs.
The GPU shape lineup is the broadest among providers evaluated: BM.GPU.H100.8 (8x NVIDIA H100 80GB, two Intel Xeon Platinum 8480+ CPUs, 2 TB DDR5, 16x 3.84 TB NVMe, 400 Gb networking, NVSwitch/NVLink 4.0 with 3.2 TB/s bisection bandwidth), BM.GPU.H200.8 (8x H200 141GB), BM.GPU.B200.8 (8x B200 192GB per NVIDIA’s standard specification [22]), BM.GPU.B300.8 (8x B300 288GB per NVIDIA’s standard specification for the Blackwell Ultra [22]), BM.GPU.MI300X.8 (8x AMD MI300X with 192 GB HBM3 per GPU), BM.GPU.MI355X.8 (8x MI355X 288GB HBM3E), BM.GPU.A100, BM.GPU.L40S.4, and BM.GPU.A10 [19]. Grace-Blackwell GB200 and GB300 superchips are also available. Networking uses RoCE v2 RDMA cluster fabric scaling to 131,072 B200 GPUs in a single Supercluster. BM.GPU.H100.8 pricing is $10 per GPU-hour ($80/hr per 8-GPU node) per Oracle’s official blog, with the same $10/GPU-hour rate confirmed for BM.GPU.H200.8 [21]. AMD MI300X is offered at $6 per GPU-hour per Oracle’s own page [20]. Data transfer is free: Oracle moved to free global egress in February 2026. Annual Universal Credits typically discount 25 to 45 percent below pay-as-you-go list pricing for GPU workloads.
Oracle’s positioning as the bare-metal hyperscaler matters concretely for two buyer profiles. For enterprise AI teams already running on OCI for non-GPU workloads, the bare-metal GPU shapes integrate with the existing identity and compliance infrastructure without requiring a separate cloud relationship. For regulated workloads requiring demonstrable hardware-level isolation under HIPAA, PCI, FedRAMP, or sector-specific requirements, the no-hypervisor architecture provides isolation guarantees that virtualized cloud cannot match. Oracle’s own marketing claims up to 20 percent better price-performance versus AWS and Azure on equivalent workloads, which is vendor-published and not independently verified.
Strengths: only hyperscaler with true bare-metal GPU instances (no hypervisor); broadest GPU shape lineup including H100, H200, B200 (192 GB), B300 (288 GB), MI300X (192 GB), MI355X (288 GB), A100, L40S, and Grace-Blackwell superchips; RoCE v2 RDMA cluster fabric scaling to 131,072 GPUs; comprehensive hyperscaler compliance (SOC 2, ISO 27001, HIPAA, PCI, FedRAMP); free global egress (since February 2026); integration with broader OCI services.
Limitations: high per-hour pricing relative to specialized AI clouds ($10/GPU-hour for H100 versus $1.99/hr at Voltage Park / Lightning AI); minimum 8-GPU shape granularity on most data-center GPUs (no single-GPU H100 BM shape); RoCE v2 RDMA less widely adopted for ML training than Quantum-2 InfiniBand; OCI ecosystem familiarity required (sales engagement typical for enterprise commitments); reserved pricing often required for Supercluster-scale capacity.
Specialized AI Bare-Metal: Voltage Park (now Lightning AI)

source: voltagepark.com
Voltage Park completed a merger with Lightning AI on January 21, 2026, and the combined company now operates under the Lightning AI name per the BusinessWire press release [24] and Cooley law firm announcement [25]. The combined fleet exceeds 35,000 owned and operated H100, B200, and GB300 GPUs across six US data centers in Texas, Virginia, Washington, and Utah [23][24], with the merged entity valued at over $2.5 billion and an annual recurring revenue exceeding $500 million per RootData coverage [26]. Voltage Park previously acquired GPU marketplace TensorDock in April 2025 to expand its catalog beyond H100 SXM5 [28].
The flagship offering remains HGX H100 dedicated reserve with single-tenant isolation, 3,200 Gbps Quantum-2 InfiniBand fabric, and dedicated clusters scaling from 64 to 4,064 HGX H100 GPUs per cluster, with on-demand pools sized from 1 to 1,016 GPUs. The hardware platform is Dell PowerEdge XE9680 servers with 8x HGX H100 SXM5 per node, 1 TB of RAM, and 4th Generation Intel Xeon Scalable processors (Sapphire Rapids, like Xeon Platinum 8480+) per Dell’s published technical guide for the XE9680 [27]. On-demand access starts at $1.99/hr per GPU with no minimum terms and self-serve provisioning in 15 minutes; long-term contracts (12+ months) are available with friendly payment terms [23]. The pricing structure has no hidden ingress, egress, or support costs. The roadmap covers H200, B200, B300, GB200, and GB300 on term contracts. Tier 3+ data centers, Palo Alto firewall, encryption, access controls, and penetration testing audits. Combined with Lightning AI’s developer platform (used by over 400,000 individual developers per the merger announcement), customers now access both the Voltage Park infrastructure and Lightning AI’s training, deployment, and inference tooling on a single platform.
For AI buyers needing single-tenant H100 SXM bare metal with InfiniBand scaling, Voltage Park / Lightning AI is positioned competitively against specialized AI clouds on price ($1.99/hr versus Lambda H100 SXM at $2.99/hr and CoreWeave H100 at $6.16/hr per the prior training article research). The constraint is GPU breadth: HGX H100 is the only currently-deployed bare-metal SKU, with newer architectures available only on long-term contracts.
Strengths: lowest published H100 bare-metal price among providers evaluated at $1.99/hr with no contracts; 3,200 Gbps Quantum-2 InfiniBand fabric scaling to 4,064 GPUs per dedicated cluster; self-serve provisioning in 15 minutes; physical isolation with dedicated bare-metal HGX H100 SXM5; no ingress, egress, or support fees; Dell PowerEdge XE9680 platform with 1 TB RAM per node and 4th Gen Intel Xeon Scalable CPUs; combined 35,000+ GPU fleet across six US data centers post-merger; integrated access to Lightning AI’s training and inference tooling.
Limitations: HGX H100 only currently deployed (H200, B200, B300, GB200, GB300 on long-term contracts only); US data centers only with no EU or APAC presence; per-hour billing rather than per-second or per-minute on on-demand tier; recent corporate consolidation following Lightning AI merger may affect product roadmap stability during integration.
Specialized AI Bare-Metal: FluidStack

source: fluidstack.io
FluidStack operates single-tenant GPU clusters built on H100, H200, B200, and GB200 hardware, with single-tenant isolation as the default deployment model. The company markets bare-metal clusters scaling from 8 to 30,000 GPUs that can be provisioned in under 48 hours, with hardware, network, and storage isolated per tenant. In November 2025, Anthropic announced a $50 billion investment in US data center infrastructure and selected FluidStack as the build partner for custom facilities in Texas and New York, with more sites to come [30]. The $50 billion figure is Anthropic’s total infrastructure spend rather than a contract value paid to FluidStack. FluidStack serves as the infrastructure operator across multiple announced sites including a 245 MW project in Louisiana (Hut 8 River Bend) [33], a 168 MW project in Texas (TeraWulf Abernathy) [31], a 244 MW project in Texas (Cipher Mining) [32], and a 360 MW project in New York (TeraWulf Lake Mariner) [30]. Google has backed approximately $1.4 billion of FluidStack’s Cipher Mining lease obligations and taken a 5.4 percent stake in Cipher [32]; Google additionally backstopped FluidStack’s TeraWulf deal in exchange for a 14 percent TeraWulf stake [31].
The FluidStack technical stack is differentiated by Atlas OS for bare-metal provisioning and Lighthouse for monitoring, which the vendor claims maintains 95 percent or higher of theoretical hardware performance with auto-restart on failures (vendor-published, not independently verified). Compliance certifications include HIPAA, GDPR, ISO 27001, and SOC 2 Type 2, with a 15-minute support SLA [29]. On-demand H100 pricing is approximately $2.10/hr per GPU with hourly billing rounding, plus reserved cluster commitments accessed through sales [29]. Multi-node InfiniBand fabrics are standard on cluster configurations.
The FluidStack positioning is enterprise: clusters at AI lab scale (the Anthropic deployment is the marquee reference) with single-tenant guarantees that hyperscaler virtualized clouds cannot match. The constraint is access: most FluidStack capacity requires sales engagement rather than self-serve provisioning, and pricing for non-standard configurations is quote-based. For buyers needing 1,000+ GPU single-tenant clusters with InfiniBand and the broadest current-generation NVIDIA hardware (B200, GB200 available), FluidStack and Voltage Park / Lightning AI are the two specialized bare-metal options that can serve frontier-scale training.
Strengths: single-tenant isolation as default deployment model; scales from 8 to 30,000 GPUs per cluster with sub-48-hour provisioning; H100, H200, B200, and GB200 available; comprehensive compliance (HIPAA, GDPR, ISO 27001, SOC 2 Type 2); 15-minute support SLA; selected by Anthropic as the build partner for its announced $50 billion US data center investment spanning multiple sites in Texas and New York; backed by Google guarantees on Cipher Mining ($1.4B) and TeraWulf deals.
Limitations: pricing largely quote-based with sales engagement required for non-standard configurations; on-demand H100 at $2.10/hr is above Voltage Park / Lightning AI’s $1.99/hr on-demand rate; multi-node InfiniBand configurations not available on all clusters by default; non-public capacity allocation favors larger customers (smaller buyers may face longer lead times); reliance on partner data centers (Cipher, TeraWulf, Hut 8) introduces dependencies on third-party operators and crypto-mining-derived infrastructure.
AMD Specialized Bare-Metal: TensorWave

source: tensorwave.com
TensorWave is the only AMD-only bare-metal GPU provider among those evaluated, building its platform exclusively on AMD Instinct accelerators (MI355X, MI325X, MI300X) with the ROCm software ecosystem. The company operates from Las Vegas with infrastructure scaling to an 8,192-GPU MI325X cluster currently being deployed. Certifications include ISO 27001, SOC 2 Type II, and HIPAA compliance.
TensorWave’s published bare-metal pricing offers MI355X starting at $2.85 per GPU-hour, MI325X starting at $1.95 per GPU-hour, and MI300X options (currently sold out per the vendor’s pricing page in May 2026) [34]. The Enterprise Cluster tier supports multi-node configurations with custom networking and storage powered by Weka file system, with pricing customized per deployment. Reservations range from 6 months to 3 years. All tiers offer guaranteed availability and full root access on bare-metal nodes; an alternative managed Kubernetes cluster option is available for orchestrated AI workloads.
For AI buyers with workloads suited to AMD’s high-VRAM accelerators (192 GB on MI300X, 256 GB on MI325X, 288 GB on MI355X), TensorWave provides a pure-play AMD bare-metal option. The MI300X family is particularly attractive for serving large language models with reduced sharding overhead: a single MI325X (256 GB) or MI355X (288 GB) holds Llama 3.1 405B at INT4 quantization (roughly 203 GB for model weights alone, per Hugging Face’s official AWQ-INT4 release [35]), where competing 80 GB or 141 GB GPUs require sharding the model across multiple cards. The constraint is the AMD ROCm ecosystem requirement: codebases that rely on custom CUDA kernels require porting to ROCm, and not all PyTorch ecosystem packages have first-class ROCm support. However, ROCm has matured significantly with first-class support in vLLM, SGLang, and PyTorch for standard transformer architectures.
Strengths: only AMD-only bare-metal provider here; high-VRAM AMD Instinct accelerators (192 GB MI300X, 256 GB MI325X, 288 GB MI355X) for serving large models without sharding; competitive pricing relative to AMD options at hyperscalers (MI325X at $1.95/hr versus Azure ND MI300X v5 at approximately $11/GPU-hr); ISO 27001, SOC 2 Type II, and HIPAA compliant; full root access on bare-metal nodes; managed Kubernetes option for orchestrated workloads.
Limitations: AMD ROCm ecosystem required (CUDA-only codebases require porting effort); single location (Las Vegas) with US-only presence; MI300X currently sold out per the vendor’s pricing page; Enterprise Cluster pricing is quote-based; ROCm support in some PyTorch ecosystem packages still maturing (verify framework compatibility before commitment); 6-month minimum reservation term on dedicated bare-metal tiers.
Common Mistakes When Choosing Bare-Metal GPU Servers for AI
The most consequential mistake is buying “bare metal” that is actually virtualized. Some providers market virtualized GPU instances under “bare metal” SKU names, often with a slightly higher price than their pure-cloud GPU SKU to suggest a premium tier. The test is whether the provider grants BIOS and firmware access, custom kernel driver installation, and direct nvidia-smi or rocm-smi without virtualization shims. If not, it is not bare metal regardless of branding. SemiAnalysis ClusterMAX evaluations have specifically flagged Crusoe as running lightweight hypervisor VMs rather than true bare metal [40].
The second mistake is sizing for peak utilization on bare metal when the workload is bursty. A team that needs H100-class compute for 4 hours per day during quarterly model retraining wastes money on a 24/7 H100 bare-metal contract. For this profile, cloud burst capacity is the right answer: a $3.99/hr H100 cloud rate at 4 hours per day works out to $479/month, well below any bare-metal H100 monthly contract. Bare metal wins on cost only above approximately 60 to 70 percent sustained utilization on H100-class hardware, with the threshold shifting lower for cheaper Ampere bare metal.
The third mistake is ignoring egress fees when comparing bare metal to cloud. Hyperscaler egress fees add 5 to 15 percent to effective cost, which often shifts the bare-metal break-even by 5 to 10 percentage points of utilization. Bare-metal providers with zero egress (Hostline, Voltage Park / Lightning AI, Cherry Servers’ 100 TB included, Oracle OCI since February 2026, FluidStack) eliminate this entirely, while providers with included egress allowances (Latitude.sh 20 TB per server, Liquid Web 10 TB per server) cap the cost predictably.
The fourth mistake is confusing a data center’s physical tier with information security certification. Tier III describes physical infrastructure redundancy under Uptime Institute standards; SOC 2 Type II and ISO 27001 describe information security controls. Hostline’s Vilnius data center is described as built to Tier III standards in the company’s published documentation, and Hostline has no publicly listed SOC 2 attestation on its site. Cherry Servers’ Lithuania GPU site holds both ISO 27001 and SOC 2 Type II per the vendor’s published location page. These certifications cover different risks, and procurement teams should verify which ones are relevant to their compliance requirements.
The fifth mistake is choosing multi-tenant cloud for regulated workloads when bare metal is the legally compliant option. Healthcare data under HIPAA, financial data under PCI DSS, defense and ITAR-controlled work, and certain EU GDPR processing scenarios require demonstrable hardware-level isolation that virtualized cloud cannot provide regardless of the provider’s compliance attestations. Oracle OCI bare-metal shapes, FluidStack single-tenant clusters, and Voltage Park / Lightning AI dedicated reserve are the three options here that combine hyperscaler-grade compliance certifications with bare-metal isolation.
Use Case Routing
For teams running LoRA fine-tuning, sustained inference serving, or RAG production workloads on 7B to 32B parameter models from the EU with GDPR data residency and a preference for predictable monthly EUR billing, the EU value tier provides the strongest fit. Hostline’s dual or triple RTX A5000 configurations (EUR 903 to 1,220 per month) deliver multi-GPU bare metal with ECC system RAM, zero egress fees, and a Vilnius data center described as built to Tier III standards in the company’s published documentation. The configuration is the only multi-GPU bare-metal option at fixed monthly EUR pricing in the EU value tier among the providers evaluated here. Hetzner GEX44 at EUR 184/month is the cheapest single-GPU entry point with native FP8 (with the caveat that system RAM is non-ECC at the entry tier), while Hetzner GEX131 at EUR 889/month provides 96 GB of GDDR7 with FP4 for higher-VRAM single-GPU inference. Cherry Servers fills the gap for buyers who need hourly billing flexibility or higher VRAM (A100 80GB at $2.18/hr, subject to pre-order availability) than Hostline’s 24 GB ceiling supports.
For full fine-tuning of 30B to 70B models with NVLink, high-VRAM single-GPU inference on 70B+ models, or production inference on H100-class hardware with US data residency, the global and US tier becomes the right fit. Liquid Web’s H100 NVL at $3.97/hr or dual H100 NVL at $6.85/hr supports 94 GB single-GPU inference with US data centers. Latitude.sh provides H100 PCIe across approximately 25 global locations for buyers needing geographic distribution (now backed by Megaport’s global private connectivity fabric post-acquisition). Oracle OCI bare-metal shapes (BM.GPU.H100.8, BM.GPU.B200.8, BM.GPU.MI300X.8) serve the regulated and hyperscaler-integrated enterprise buyer at approximately $10/GPU-hour.
For multi-node distributed training and frontier-scale AI workloads requiring 1,000 or more GPUs with InfiniBand fabric, the specialized AI tier provides the only viable bare-metal options. Voltage Park / Lightning AI HGX H100 dedicated reserve at $1.99/hr on-demand or long-term contracts scales to 1,016 GPUs on the on-demand tier and 64 to 4,064 GPUs per cluster on the dedicated reserve tier, with 3,200 Gbps Quantum-2 InfiniBand and access to Lightning AI’s developer tooling post-merger. FluidStack single-tenant clusters from 8 to 30,000 GPUs (with the Anthropic partnership as the marquee reference, FluidStack being the build partner for Anthropic’s announced $50 billion US data center investment) cover the full range of frontier-scale workloads on H100, H200, B200, and GB200 hardware. For AMD-specific workloads (vLLM ROCm, MI325X or MI355X inference on Llama 3.1 405B with reduced sharding), TensorWave is the only AMD-only bare-metal specialist with MI355X at $2.85/GPU-hour, MI325X at $1.95/GPU-hour, and MI300X options.
Three Findings
The definitional gate matters more than the price comparison. Half the “best bare-metal GPU server” listicles in current SERPs fail at the definitional test by including providers that are not actually bare metal. Apply the no-hypervisor, single-tenant, direct-hardware-access test before any pricing comparison. The three categories of bare-metal marketing (true bare metal, hypervisor-on-bare-metal, and virtualized cloud sold as bare metal) produce different performance characteristics, different cost structures, and different compliance properties.
EU teams have stronger bare-metal options than the US-centric narrative typically presents. Hostline, Hetzner, and Cherry Servers provide a viable EU value tier alternative to US-centric specialized AI clouds. For LoRA fine-tuning, sustained inference, and RAG workloads on models up to 32B parameters, the EU value tier consistently beats US cloud on cost above 60 to 70 percent utilization (or lower for cheaper Ampere bare metal) while delivering GDPR data residency that US clouds cannot match. Hostline provides the lowest sustained monthly cost for multi-GPU bare metal evaluated here at EUR 903/month for dual A5000 (~$985 at May 2026 exchange rates), the only multi-GPU-in-one-chassis configuration at fixed monthly EUR pricing among the EU value tier providers evaluated, and ECC DDR4 system RAM at the entry-level GPU plan (a distinction from Hetzner GEX44 at the same EU value tier price band, where the i5-13500 consumer Raptor Lake CPU ships non-ECC DDR4 system RAM). Hetzner provides the highest single-GPU VRAM in the EU value tier at 96 GB on the GEX131. Cherry Servers provides the broadest GPU menu and hourly billing flexibility from the same Lithuania region.
The right bare-metal provider depends on workload pattern as much as on GPU generation. A 24/7 inference workload on a 7B model is better served by EUR 360/month Ampere bare metal (Hostline single A4000) than by $3.99/hr Hopper cloud, despite the older GPU architecture. A 1,000-GPU frontier pretraining run is better served by FluidStack or Voltage Park / Lightning AI InfiniBand clusters than by any single-node bare-metal provider. The framework, not any single provider’s positioning, is what the reader should carry away from this comparison.
FAQ
What is the difference between “bare-metal GPU” and “dedicated GPU instance” in cloud marketing?
A true bare-metal GPU server gives the customer direct physical access to the hardware with no hypervisor layer, including BIOS access, custom kernel drivers, and direct nvidia-smi or rocm-smi. A “dedicated GPU instance” typically runs a hypervisor (KVM, Hyper-V, or proprietary) that presents the GPU to the customer’s VM through PCIe passthrough. The dedicated instance is a single-tenant VM, but the hypervisor adds 2 to 10 percent overhead on typical AI workloads (up to 25 percent on memory-bandwidth-bound inference per Hivelocity and Aethir analyses) and restricts BIOS-level configuration. Hostline, Hetzner GEX, Cherry Servers, Voltage Park / Lightning AI, FluidStack, TensorWave, Oracle OCI BM shapes, Liquid Web Metal, and Latitude.sh Metal are true bare metal. Most “dedicated GPU instances” at hyperscalers and neoclouds are not.
Which bare-metal GPU server in this comparison is cheapest for running LoRA fine-tuning on a 13B model?
Hetzner GEX44 at EUR 184/month is the lowest absolute monthly price with 20 GB GDDR6 ECC VRAM and native FP8 support, sufficient for LoRA fine-tuning of 7B to 14B models at FP16 or quantized 30B+ models at INT4. Hostline single RTX A4000 at EUR 360/month adds 64 GB DDR4 ECC system RAM (versus GEX44’s 64 GB DDR4 non-ECC, since the i5-13500 has no system ECC support) and iDRAC 9 Enterprise remote management for production reliability. Both serve the LoRA workload at lower cost than any cloud H100 alternative running 24/7. For 13B models specifically, the GEX44’s 20 GB and the A4000’s 16 GB both work at LoRA precision; full fine-tuning of 13B at FP16 requires 24 GB+, fitting on Hostline’s dual A5000 at EUR 903/month.
Can I run distributed multi-node training on bare-metal GPU servers?
Only on providers with InfiniBand or equivalent RDMA fabric. The bare-metal providers in this comparison split cleanly into single-node-only and multi-node-capable. Single-node-only: Hostline (1 Gbps Ethernet), Hetzner GEX (1 Gbps Ethernet, single-GPU), Cherry Servers (up to 10 Gbps Ethernet), Liquid Web (10 Gbps Ethernet), Latitude.sh (no InfiniBand on multi-GPU). Multi-node-capable with InfiniBand or RDMA: Voltage Park / Lightning AI (3,200 Gbps Quantum-2 InfiniBand), FluidStack (InfiniBand on cluster configurations), Oracle OCI (RoCE v2 RDMA scaling to 131,072 GPUs). For distributed training of 70B+ models requiring more than 8 GPUs in a single training job, only these three providers can serve the workload at bare metal.
Why do some “bare-metal” providers offer hourly billing if bare metal requires physical hardware provisioning?
The standard model is that the provider pre-stages physical servers in racks and provisions them to customers using firmware-level imaging tools (typically iPXE plus a customer-selected OS image). Provisioning time is 3 to 10 minutes for providers with mature automation (Latitude.sh, Voltage Park / Lightning AI) and 1 business day for providers with less automation (Hostline, Hetzner). The hourly billing reflects the time the server is allocated to the customer, not the time spent on physical hardware moves. The customer gets exclusive physical access for the duration of the allocation. Voltage Park / Lightning AI’s 15-minute provisioning, Latitude.sh’s 3 to 10 minute provisioning, and Liquid Web’s portal provisioning all use this approach.
What certifications should I look for on a bare-metal GPU server for regulated AI workloads?
The certification stack depends on the regulatory regime. For healthcare data under HIPAA, look for HIPAA compliance attestations (Oracle OCI, FluidStack, TensorWave, Liquid Web). For financial data under PCI DSS, look for PCI compliance (Oracle OCI, hyperscaler bare-metal). For EU data residency under GDPR, look for EU data centers plus GDPR-aligned certifications (Hostline, Hetzner, Cherry Servers). For general information security baselines, look for ISO 27001 and SOC 2 Type II (most providers evaluated here hold at least ISO 27001). For physical data center reliability, look for Uptime Institute Tier III or Tier IV. These are different certifications covering different risks. Hostline’s data center in Vilnius is described as built to Tier III standards in the company’s published documentation but has no publicly listed SOC 2 attestation on its site, while Cherry Servers’ Lithuania GPU site holds both ISO 27001 and SOC 2 Type II per the vendor’s published location page.
How do I measure the actual performance difference between bare metal and virtualized cloud for my specific AI workload?
Run the same workload on both for 48 hours, measure three things, and compare. First, measure tokens per second at the target batch size and precision on each platform using the same model and framework version. Second, measure GPU memory utilization percentage at steady state, time-to-first-token at the 95th percentile for inference workloads, and step time variance for training workloads. Third, measure end-to-end cost including egress fees and any cloud-specific overhead. The performance gap is workload-dependent: compute-bound large-batch training shows the smallest gap (often under 5 percent), memory-bandwidth-bound inference shows the largest (up to 25 percent in poorly tuned environments per Hivelocity and Aethir analyses, typically 5 to 15 percent in well-tuned environments). The cost gap depends on utilization: above 60 to 70 percent utilization on H100-class hardware (lower thresholds on cheaper Ampere bare metal), bare metal typically wins regardless of the performance gap. Below 40 percent utilization, cloud wins because idle time costs nothing.
References
[1] vLLM Project. “Distributed Inference and Serving.” vllm.ai/docs. Accessed May 2026. Supports tensor parallelism requirement that tensor_parallel_size evenly divides both query and KV head counts (Llama-class 70B: 64 query heads, 8 KV heads).
[2] NVIDIA Corporation. “RTX A5000 Product Page.” nvidia.com/en-us/design-visualization/rtx-a5000. Accessed May 2026. Supports RTX A5000 NVLink bridge documented as 2-way only between two cards.
[3] Hostline UAB. “GPU Dedicated Servers.” hostline.io/dedicated-servers/gpu-servers. Accessed May 2026. Supports Hostline pricing (EUR 360 / EUR 903 / EUR 1,220 per month), Intel Xeon Gold 6130 platform, RTX A4000 16 GB and RTX A5000 24 GB configurations, DDR4 ECC system RAM at all tiers, iDRAC 9 Enterprise on all plans, no setup fee, zero egress fees, DDoS protection, N+1 redundancy claim on power and networking, 1 Gbps network port. Owned and operated by HOSTLINE UAB (Lithuanian company registration number 302660481).
[4] Hetzner Online GmbH. “Dedicated Root Server Matrix – GPU.” hetzner.com/dedicated-rootserver/matrix-gpu. Accessed May 2026. Supports GEX44 at EUR 184/month with EUR 79 setup fee (RTX 4000 SFF Ada 20 GB GDDR6 ECC VRAM, Intel Core i5-13500 consumer Raptor Lake, 64 GB DDR4 non-ECC system RAM,2x 1.92 TB NVMe Gen3 RAID 1); GEX131 at EUR 889/month monthly commitment or EUR 1.42/hr on the separate on-demand tier (RTX PRO 6000 Blackwell Max-Q 96 GB GDDR7 ECC, 256 GB DDR5 ECC expandable to 768 GB, 2x 960 GB NVMe Gen4). Hetzner statement that GEX servers are single-GPU only. Falkenstein and Nuremberg data centers, ISO 27001 certification, 100 percent green energy.
[5] Puget Systems. “RTX 6000 Ada Generation Max-Q Performance Comparison.” pugetsystems.com. Accessed May 2026. Supports Max-Q variant performance approximately 5 to 14 percent below the 600W Workstation Edition while drawing 300W.
[6] Cherry Servers UAB. “AI / GPU Servers.” cherryservers.com/ai-servers. Accessed May 2026. Supports A100 80GB at $2.18/hr or $1,530.17/month (pre-order), A40 48GB at $0.74/hr or $436.30/month (pre-order), A16 64GB at $0.498/hr or $290.52/month (waitlist), A10 24GB at $0.479/hr or $279.72/month (waitlist), A2 16GB at $0.22/hr or $128.52/month, Tesla P4 at $0.157/hr; Custom Dedicated Server base from $158.13/month; AMD Ryzen 7700X, AMD EPYC 7402P, and Intel Xeon Gold 6230R CPU options; EPYC base supports up to 2 GPUs per chassis; up to 1024 GB ECC DDR4 RAM and 80 TB storage; 100 TB/month free egress with overage at approximately EUR 0.5/TB.
[7] Cherry Servers UAB. “Data Center Locations.” cherryservers.com/network/locations. Accessed May 2026. Supports Lithuania (Šiauliai) GPU site certifications: ISO 27001, ISO 22301, ISO 9001, SOC 1 Type II, and SOC 2 Type II.
[8] Cherry Servers UAB. “Cherry Servers Tokyo Now Live.” cherryservers.com/blog/cherry-servers-tokyo-location-live. Accessed May 2026. Supports Tokyo location going live in May 2026; combined 7 global data center locations (Lithuania, Amsterdam, Frankfurt, Stockholm, Chicago, Singapore, Tokyo).
[9] Cherry Servers UAB. “When to Use Bare Metal Servers: Complete Guide.” cherryservers.com/blog/when-to-use-bare-metal-servers. Accessed May 2026. Supports bare-metal break-even threshold around 70 percent utilization with payback typically within 6 to 12 months.
[10] Latitude.sh. “Metal GPU Dedicated Clusters.” latitude.sh. Accessed May 2026. Supports NVIDIA H100 PCIe, RTX 6000 Ada, RTX PRO 6000, and L40S availability; dual 100 Gbps networking on multi-GPU nodes; 20 TB free outbound bandwidth per server per month; bare-metal provisioning in 3 to 10 minutes; approximately 25 global locations across North America, Europe, Latin America, and Asia-Pacific including Sydney added in 2025.
[11] Latitude.sh. “How Bare Metal Servers Can Nullify Outages.” latitude.sh/blog. Accessed May 2026. Supports service level agreements of up to 99.9 percent.
[12] Latitude.sh. “Pricing.” latitude.sh/pricing. Accessed June 2026. Supports Latitude.sh published hourly and monthly GPU bare-metal pricing across its global locations.
[13] Atlantic.net. “Latitude.sh Monthly Pricing Reference.” atlantic.net. Accessed May 2026. Supports single H100 PCIe server at approximately $2,995.20/month and single A100 80GB server at approximately $2,234.88/month on fixed monthly billing.
[14] Spheron Network. “Bare Metal Provider Analysis.” spheron.network. Accessed May 2026. Supports Latitude.sh 3 to 10 minute provisioning timeframe and H100 regional availability (Dallas, São Paulo, Tokyo).
[15] PitchBook Data, Inc. “Latitude.sh Company Profile.” pitchbook.com/profiles/company/490846-60. Accessed May 2026. Supports Latitude.sh São Paulo, Brazil headquarters and acquisition by Megaport Limited (ASX: MP1) on November 27, 2025.
[16] Megaport Limited. “Megaport to Acquire Latitude.sh.” megaport.com investor announcement. November 2025. Accessed May 2026. Supports the Megaport (ASX: MP1) acquisition of Latitude.sh and integration of Latitude.sh’s compute platform with Megaport’s global private connectivity fabric.
[17] Liquid Web. “GPU Hosting.” liquidweb.com/gpu-hosting. Accessed May 2026. Supports NVIDIA L4 Ada 24 GB at $1.07/hr (dual AMD EPYC 9124, 128 GB DDR5, 1.92 TB NVMe), L40S 48 GB at $1.92/hr, H100 NVL 94 GB at $3.97/hr, 2x H100 NVL at $6.85/hr; 10 TB outbound bandwidth per server; US data centers documented in Lansing, Michigan and Phoenix, Arizona.
[18] PR Newswire. “Liquid Web Launches GPU Metal Bare-Metal Servers.” prnewswire.com. October 2024. Accessed May 2026. Supports the vendor claim of “up to 15% better GPU performance over virtualized environments.”
[19] Oracle Corporation. “Compute GPU Shapes.” oracle.com/cloud/compute/gpu. Accessed May 2026. Supports BM.GPU.H100.8 (8x H100 80 GB, 2x Xeon Platinum 8480+, 2 TB DDR5, 16x 3.84 TB NVMe, 400 Gb networking, NVSwitch/NVLink 4.0 with 3.2 TB/s bisection bandwidth), BM.GPU.H200.8 (8x H200 141 GB), BM.GPU.B200.8, BM.GPU.B300.8, BM.GPU.MI300X.8 (192 GB HBM3 per GPU), BM.GPU.MI355X.8 (288 GB HBM3E per GPU), BM.GPU.A100, BM.GPU.L40S.4, and BM.GPU.A10; RoCE v2 RDMA cluster fabric scaling to 131,072 B200 GPUs; SOC 2, ISO 27001, HIPAA, PCI, and FedRAMP compliance.
[20] Oracle Corporation. “Compute Pricing.” oracle.com/cloud/compute/pricing. Accessed May 2026. Supports AMD MI300X at $6/GPU-hour; free global egress (since February 2026); Annual Universal Credits reducing list rates by 25 to 45 percent for GPU workloads.
[21] Oracle Corporation. “Now Generally Available: The Largest, Fastest AI Supercomputer in the Cloud.” blogs.oracle.com/cloud-infrastructure. Accessed May 2026. Supports BM.GPU.H100.8 list price at $10 per GPU-hour ($80/hr per 8-GPU node) and the same $10 per GPU-hour rate confirmed for BM.GPU.H200.8.
[22] NVIDIA Corporation. “NVIDIA Blackwell Architecture Datasheet.” nvidia.com/blackwell. Accessed May 2026. Supports B200 standard specification of 192 GB HBM3e and B300 (Blackwell Ultra) standard specification of 288 GB HBM3e.
[23] Voltage Park. “Pricing.” voltagepark.com/pricing. Accessed May 2026. Supports on-demand HGX H100 at $1.99 per GPU-hour, on-demand pool sized 1 to 1,016 GPUs, dedicated reserve clusters 64 to 4,064 GPUs, 3,200 Gbps Quantum-2 InfiniBand fabric, 15-minute self-serve provisioning, and six US data centers in Texas, Virginia, Washington, and Utah.
[24] BusinessWire. “Lightning AI and Voltage Park Complete Merger to Create the First Cloud Built for AI.” businesswire.com. January 21, 2026. Accessed May 2026. Supports the merger completion date, the combined company operating under the Lightning AI name, the combined fleet of 35,000+ owned and operated H100, B200, and GB300 GPUs, and Lightning AI’s 400,000+ developer user base pre-merger.
[25] Cooley LLP. “Lightning AI and Voltage Park Complete Merger.” cooley.com/news/coverage/2026/2026-01-21-lightning-ai-and-voltage-park-complete-merger. January 21, 2026. Accessed May 2026. Supports legal counsel confirmation of merger completion under the Lightning AI name.
[26] RootData. “Lightning AI Merges with Voltage Park, Valued at Over $2.5 Billion.” rootdata.com/news/513884. January 2026. Accessed May 2026. Supports combined valuation exceeding $2.5 billion and annual recurring revenue exceeding $500 million.
[27] Dell Technologies. “Dell PowerEdge XE9680 Technical Guide.” delltechnologies.com/asset/en-ca/products/servers/technical-support/poweredge-xe9680-technical-guide.pdf. Accessed May 2026. Supports XE9680 use of 4th Generation Intel Xeon Scalable processors (Sapphire Rapids, up to 56 cores per CPU) or 5th Generation (up to 64 cores per CPU); 32 DDR5 DIMM slots; 8x NVIDIA HGX H100 80GB SXM5 GPU configuration; 1 TB to 4 TB RAM capacity.
[28] DataCenterDynamics. “GPUaaS Provider Voltage Park Merges with Cloud Platform Lightning AI.” datacenterdynamics.com/en/news/gpuaas-provider-voltage-park-merges-with-cloud-platform-lightning-ai. Accessed May 2026. Supports the Voltage Park acquisition of TensorDock in April 2025 and the 24,000 H100 fleet acquisition.
[29] FluidStack. “Single-Tenant Bare-Metal Clusters.” fluidstack.io. Accessed May 2026. Supports H100, H200, B200, and GB200 NVL72 configurations; Atlas OS bare-metal provisioning; Lighthouse monitoring with vendor-claimed 95 percent or higher of theoretical hardware performance; HIPAA, GDPR, ISO 27001, and SOC 2 Type 2 compliance; 15-minute support SLA; on-demand H100 at approximately $2.10 per GPU-hour.
[30] DataCenterDynamics. “Anthropic Plans $50bn US Data Center Spend, Starting with Fluidstack Sites in Texas and New York.” datacenterdynamics.com/en/news/anthropic-plans-50bn-us-data-center-spend-starting-with-fluidstack-sites-in-texas-and-new-york. Accessed June 2026. Supports Anthropic’s announced $50 billion US data center investment (announced November 2025), FluidStack selected as the build partner for custom facilities in Texas and New York, and the multi-site deployment across Hut 8 (River Bend Louisiana 245 MW), TeraWulf (Abernathy Texas and Lake Mariner New York), and Cipher Mining (Texas).
[31] Sacra. “Fluidstack Revenue, Funding and Growth Rate.” sacra.com/c/fluidstack. April 2026. Accessed May 2026. Supports TeraWulf 168 MW Abernathy joint venture; Google approximately $1.3 billion TeraWulf-related lease backing; Cipher Mining 207 MW Texas partnership scope.
[32] Alpha Spread. “Cipher Mining Secures $3 Billion AI Data Hosting Deal with Fluidstack; Google Takes Stake.” alphaspread.com. September 2025. Accessed May 2026. Supports Google backstop of approximately $1.4 billion of FluidStack’s Cipher Mining lease obligations and Google’s approximately 5.4 percent equity stake in Cipher Mining.
[33] Yahoo Finance. “Hut 8 Shares Soar as Data Center Firm Inks $7 Billion Revenue Deal with Fluidstack, Anthropic.” finance.yahoo.com. December 17, 2025. Accessed May 2026. Supports the Hut 8 245 MW Louisiana River Bend partnership, $7 billion 15-year base contract value (up to $17.7 billion with extensions), and Google financial backstop.
[34] TensorWave. “Bare Metal AMD MI Series.” tensorwave.com/bare-metal. Accessed May 2026. Supports MI355X starting at $2.85 per GPU-hour, MI325X starting at $1.95 per GPU-hour, MI300X currently sold out, Enterprise Cluster tier with Weka file system, 6-month to 3-year reservation terms, ISO 27001 / SOC 2 Type II / HIPAA compliance, Las Vegas data center, and 8,192-GPU MI325X cluster currently deploying.
[35] Hugging Face. “Llama 3.1 405B AWQ-INT4 Release Notes.” huggingface.co. Accessed May 2026. Supports Llama 3.1 405B at INT4 quantization weight size of approximately 203 GB.
[36] Principled Technologies. “NVIDIA A100 Virtualized vs Bare-Metal ML Training Benchmark Report.” principledtechnologies.com. Accessed May 2026. Supports virtualized A100 reaching approximately 97.5 percent of bare-metal performance (approximately 2.5 percent overhead) on standard ML training benchmarks.
[37] Springer. “GPU Container and VM Performance Benchmark Study.” link.springer.com. Accessed May 2026. Supports approximately 5 percent container overhead, approximately 10 percent VM overhead, and approximately 15 percent combined overhead.
[38] Hivelocity. “Virtualization Tax: Bare Metal vs Virtualized GPU.” hivelocity.net. Accessed May 2026. Supports characterization of virtualization tax at 5 to 25 percent.
[39] Aethir. “Real-World GPU Virtualization Overhead Analysis.” aethir.com. Accessed May 2026. Supports real-world deployment overhead of 15 to 25 percent.
[40] SemiAnalysis. “ClusterMAX Framework: Neocloud Classification.” semianalysis.com. Accessed May 2026. Supports classification of Crusoe as running lightweight hypervisor VMs rather than true bare metal.
Editorial Note
This article is published on hostline.io by Hostline. Hostline is one of the providers compared and is positioned for EU bare-metal AI workloads at fixed monthly EUR pricing, appearing at position 2 within the EU value tier. The same evaluation template applies to every provider. Providers are presented in narrative ordering within each tier rather than ranked by a single metric; the precise per-attribute comparison appears in the comparison table. No provider is ranked #1 overall or described as “best.”
Where competitors outperform Hostline on specific dimensions, this article states so directly. Hetzner GEX44 offers lower entry pricing (EUR 184 versus EUR 360) with native FP8 tensor core support that Hostline’s Ampere hardware does not provide, though Hetzner GEX44 uses non-ECC consumer DDR4 system RAM at the entry tier where Hostline provides ECC system RAM. Hetzner GEX131 offers 4x the VRAM per GPU (96 GB versus 24 GB) with Blackwell FP4 and FP8 support and NVMe Gen4 storage. Cherry Servers offers a broader GPU menu (A100, A40, A16, A10, A2), hourly billing flexibility, and 100 TB/month free egress. Latitude.sh offers approximately 25 global locations with API-first provisioning. Liquid Web offers H100 NVL 94 GB single-GPU configurations from US data centers. Oracle OCI offers the broadest GPU shape lineup including H100, H200, B200 (192 GB), B300 (288 GB), MI300X (192 GB), MI355X (288 GB), and Grace-Blackwell superchips with the only true bare-metal hyperscaler experience. Voltage Park / Lightning AI offers the lowest published H100 bare-metal rate at $1.99/hr with 3,200 Gbps Quantum-2 InfiniBand scaling to 4,064 GPUs per dedicated cluster, now integrated with Lightning AI’s developer platform post-merger. FluidStack offers single-tenant clusters from 8 to 30,000 GPUs, serving as the build partner for Anthropic’s announced $50 billion US data center investment. TensorWave offers the only AMD-only bare-metal option with MI300X family accelerators up to 288 GB HBM3E.
Agneta Venckutė