The GPU cloud for your next leap

Your ambition.Our power.AI at scale.

Train more ambitious models. Bring your AI to production. With dedicated NVIDIA GPUs, high-speed networking, and a cluster built around your business.

Start with 8 GPUs from $19.50/hr · yearly reservation

Electricity, human creativity, and cooling connected to a central NVIDIA GPU cluster, with AI tokens emerging from computation.
  • Qwen
  • DeepSeek
  • Llama
  • Kimi K3
Dedicated power for your AINVIDIA DGX VR200
Hypervisor overhead
0%
InfiniBand per node
3.2 Tbps
Cluster goodput, 90 days
97.4%
Power across 5 sites
1 GW

Behind the teams turning AI into business impact.

  • Bank of America
  • Halliburton
  • BRF
  • PicPay
  • BMG Bank
  • City of Rio de Janeiro
  • CNSesi
  • AICUBE
  • AirCloud

01Find your compute

Your next AI breakthrough starts with the right GPU.

Training, fine-tuning, or inference: every workload has its own demands. Find the capacity that fits your team, your model, and your budget.

See all pricing
NVIDIA HGX

B300

Memory-rich compute for large models and inference

288 GB HBM3eMemory per GPU
Inside the node
8 GPU · 30 TB NVMe · 2 TB RAM
Node interconnect
Quantum-X800 XDR 800G
Explore the configuration

GPU interconnectNVLink 5 · 1.8 TB/s per GPU

Blackwell Ultra

Explore HGX B300 details ↗

Yearly · reserved

from$58.50/node-hr

Get a quote
NVIDIA HGX

H200

More memory for long-context model serving

141 GB HBM3eMemory per GPU
Inside the node
8 GPU · 30 TB NVMe · 2 TB RAM
Node interconnect
8×400G IB NDR · 3.2 Tbps
Explore the configuration

GPU interconnectNVLink 4 · 900 GB/s

Explore HGX H200 details ↗

Yearly · reserved

from$31.20/node-hr

Get a quote
NVIDIA HGX

H100

Training and fine-tuning with your budget in mind

80 GB HBM3Memory per GPU
Inside the node
8 GPU · 15 TB NVMe · 2 TB RAM
Node interconnect
8×400G IB NDR · 3.2 Tbps
Explore the configuration

GPU interconnectNVLink 4 · 900 GB/s

Workhorse

Explore HGX H100 details ↗

Yearly · reserved

from$19.50/node-hr

Get a quote

Prices in USD per 8-GPU node-hour. Rack-scale systems are quoted separately. Subject to capacity and contract terms.

06From research to results

Your next discovery. Your next product. Your next scale.

  • Train what comes next

    Give your foundation models the capacity they demand. Reserved clusters, dedicated networking, and spare nodes support the infrastructure behind your most ambitious training runs.

    Explore B200, GB200, and GB300 →
  • Turn knowledge into an advantage

    Adapt models to your data and business challenges. Dedicated 8-GPU nodes and persistent storage give your team room to test, fine-tune, and evolve.

    Find your H100 or H200 →
  • Put your AI to work

    Bring large models to production with memory-rich GPUs. Serve long-context applications on a dedicated infrastructure foundation you can plan around.

    Discover H200 and B300 →

02The TensorBay advantage

You bring the talent. We bring the foundation.

Every layer of your cluster is designed to give your team more control, more predictability, and more room to build.

  • F-01

    Every GPU works for your model

    Run directly on dedicated hardware, with the machine’s capacity focused on your workload. Take control of performance at every stage of your project.

  • F-02

    A cluster with your name on it

    Dedicated nodes, networking, and storage give your team an exclusive foundation to train, experiment, and put AI to work.

  • F-03

    Confidence before the first run

    Your cluster goes through 72 hours of load, memory, network, and thermal testing. Review the results and approve delivery with the evidence in hand.

  • F-04

    The freedom to build your way

    Your drivers, your kernels, your tools. Work with Slurm, Kubernetes, or SSH and shape the environment around your model’s needs.

  • F-05

    Continuity built into the plan

    Spare nodes in the same rack and a 15-minute replacement SLA help reduce the impact of hardware failures on your project.

  • F-06

    Clarity for your next decision

    Published rates, per-minute metering within your reservation, and 20 TB of outbound data included per node each month. Plan with the costs in view.

03Infrastructure behind your ambition

Every breakthrough needs a foundation to match.

Power, cooling, GPUs, and networking. We design and operate every layer together to give your AI a foundation ready to grow.

  1. 01Power

    1 GW

    Contracted power capacity across five countries to support the expansion of our infrastructure.

  2. 02Cooling

    PUE 1.12

    Direct-to-chip liquid cooling in 130 kW racks. Engineered for the demands of intensive AI workloads.

  3. 03GPU capacity

    24,000+ GPUs

    Capacity expands in six-week cycles, with 72 hours of testing before each handover.

  4. 04Useful work

    97.4%

    Cluster goodput over the last 90 days: the share of capacity translated into useful work.

04Networking and storage

Your GPUs. Working as one.

Bring compute, data, and storage together in infrastructure dedicated to your model, from the first training batch to the final checkpoint.

Plan my cluster

Dedicated networking

A network built for your GPUs.

Dedicated InfiniBand connects your cluster with bandwidth for distributed training and data exchange between GPUs.

800Gb/s

per InfiniBand XDR link, Quantum-X800 for Blackwell

Network specifications
Training fabric — Hopper
8×400 Gb/s InfiniBand NDR per node (3.2 Tbps), rail-optimized non-blocking fat-tree, SHARP in-network reduction
Training fabric — Blackwell
Quantum-X800 InfiniBand XDR, 800G per link; GB200 racks add a 72-GPU coherent NVLink 5 domain at 1.8 TB/s per GPU
Oversubscription
1:1, all tiers. No blended east-west fabric, no shared spine with other tenants
Front-end network
Dual 100 GbE per node, DDoS-protected, BYO-IP supported

Storage

Keep your data in stride.

Local NVMe, WEKA, and S3-compatible storage to load datasets and save checkpoints at the scale your project needs.

Up to

720GB/s

aggregate read throughput per pod, with dedicated WEKA

Storage specifications
Local scratch
30 TB NVMe Gen5 per node, ~55 GB/s read — checkpoint staging without touching the network
Parallel filesystem
Managed WEKA, dedicated per cluster: up to 720 GB/s aggregate read per pod, POSIX + GPUDirect Storage
Object storage
S3-compatible, NVMe-cached, co-located with compute — $0.055/GB-month hot tier

Data transfer

Take your data further.

Free ingress and traffic between nodes, with an included egress allowance to help you plan your operation.

20TB

egress included per node, per month

Data transfer terms
Data transfer
Ingress $0 · 20 TB/node-month egress included, then $1.25/TB · inter-node $0. Free 100G Direct Connect on yearly reservations

Networking and storage are configured for your cluster architecture.View rates and terms

05Invest in your next move

Compute to grow. Clarity to invest.

Choose your reservation term, compare capacity, and plan your investment. Here, the conversation starts with open pricing.

USD per hour, complete 8-GPU node

GPU nodes, USD per hour, complete 8-GPU node
GPU and memoryMonthly termCapacity reserved on a monthly termYearly termLower hourly rate
NVIDIAHGX B300288 GB HBM3e per GPU
Monthly
from
US$ 83.20/h
Quote monthly
Yearly
from
US$ 58.50/h
Quote yearly
NVIDIAHGX B200180 GB HBM3e per GPU
Monthly
from
US$ 72.80/h
Quote monthly
Yearly
from
US$ 50.70/h
Quote yearly
NVIDIAHGX H200141 GB HBM3e per GPU
Monthly
from
US$ 44.20/h
Quote monthly
Yearly
from
US$ 31.20/h
Quote yearly
NVIDIAHGX H10080 GB HBM3 per GPU
Monthly
from
US$ 28.60/h
Quote monthly
Yearly
from
US$ 19.50/h
Quote yearly

The hourly rate covers the complete 8-GPU node. Choose a term to request your quote.

SLA
99.9% monthly + 15-min node replacement + goodput credits
Scale
128–10,000+ GPUs · dedicated fabric per tenant
Billing
Per-minute metering inside the reservation · $0 ingress

Every detail accounted for

Your infrastructure. No blind spots.

See what is included and explore rates for storage, data transfer, and networking.

Talk to a specialist
Storage
  • Local NVMe scratch — 30 TB per nodeIncluded
  • Parallel filesystem — managed WEKAdedicated per cluster$0.09/GB-mo
  • Object storage — hot, S3-compatibleNVMe-cached, co-located$0.055/GB-mo
  • Object storage — archive$0.015/GB-mo
  • Snapshots$0.021/GB-mo
Data transfer
  • Ingressalways$0
  • Egress included20 TB/node-mo
  • Egress overageall sites$1.25/TB
  • Inter-node / east-westInfiniBand fabric$0
Networking
  • East-west InfiniBand fabricdedicated per clusterIncluded
  • Private VLAN / VPCIncluded
  • Additional public IPv4IPv6 free$4/IP-mo
  • BYO-IP — /24 v4 or /48 v6$200 setupFree
  • Managed firewall$4/node-mo
  • DDoS protectionIncluded
  • Direct Connect 10G / 100G100G free on yearly terms$450 / $1,800-mo

USD prices per 8-GPU HGX node-hour; rack-scale systems are quoted per rack. Taxes excluded. Managed Slurm and Kubernetes control planes are included. Storage and Direct Connect follow the rates above. Ask about long-term contract options.

Get my proposal

07Confidence to move forward

Your innovation deserves a secure foundation.

Bring projects to production with dedicated hardware, physical isolation, and access controls. Security is part of the infrastructure, from the first node.

  • SOC 2 Type II

    Scope: bare-metal compute, fabric, and managed storage

  • ISO 27001:2022

    Certified — all datacenter sites

  • HIPAA

    BAA available on dedicated clusters

  • GDPR

    EU region with in-country data residency

  • Hardware, networking, and storage physically dedicated to your operation
  • Encryption at rest and secure data erasure when hardware is released
  • Certified destruction of retired storage media
  • Enterprise sign-in, hardware-key authentication, and immutable audit logs
  • Review scope and reports at trust.tensorbay.com

Frequently asked questions

Your next cluster starts with clarity.

Talk to a specialist
Is GPU pricing per GPU or per node?

Published HGX prices are in USD per hour for a complete node with 8 GPUs, within a monthly or yearly reservation. NVL72 systems are quoted per rack. Taxes are not included.

Which NVIDIA GPU should I choose for my AI project?

Compare GPU memory, network requirements and your training or inference workload. The hardware section lists H100, H200, B200, B300 and full rack systems. Our team can help size a configuration based on your model, dataset and capacity requirements.

Can I use a dedicated cluster for training, fine-tuning and inference?

Yes. TensorBay offers dedicated NVIDIA GPU infrastructure for model training, fine-tuning and inference. You can work with Slurm, Kubernetes or SSH and configure the software environment for your workload.

What happens after I request a quote?

Share the GPU, capacity and reservation term you need, or describe your project. Our team reviews the requirements and prepares a proposal within 48 hours, including configuration and delivery timing. Capacity and final terms are confirmed in the proposal.

08Give your idea room to grow

Your next AI breakthrough starts with a conversation.

Bring us your challenge. Our team will help define your GPUs, capacity, and timeline. Get a proposal within 48 hours to launch your project or take it further.

  • A configuration built around your project
  • A proposal with a delivery timeline
  • Acceptance testing before you start
tensorbay@tensorbay.io