Live GPU
  • L40S 48GB from $1.29/hr in New York
  • NVIDIA A100 from $4.36/hr in London
  • NVIDIA A100 PCIe from $1.09/hr in London
  • NVIDIA A100 SXM from $1.59/hr
  • NVIDIA A16 from $0.24/hr in New Jersey
  • NVIDIA A40 from $0.29/hr in New Jersey
  • NVIDIA B200 SXM from $5.20/hr in Helsinki
  • NVIDIA B300 SXM from $7.89/hr
  • NVIDIA H100 PCIe from $2.20/hr in Montreal
  • NVIDIA H100 SXM from $1.89/hr in Helsinki
  • NVIDIA H200 NVL from $3.79/hr
  • NVIDIA H200 SXM from $4.59/hr
  • NVIDIA L4 from $0.49/hr
  • NVIDIA L40 from $0.88/hr in Montreal
  • NVIDIA L40S from $1.27/hr in Helsinki
  • NVIDIA RTX 5000 from $0.85/hr in Gravelines
  • NVIDIA RTX PRO 6000 Blackwell Server Edition from $2.09/hr
  • NVIDIA RTX PRO 6000 Blackwell Server Edition MIG 1g.24gb from $0.59/hr
  • Radeon RX 6700 XT from $0.21/hr

Catalog / GPU Stock

GPU STOCK

Canonical NVIDIA, AMD, and Intel GPUs available through stack8s — one model name each, with VRAM, architecture, bandwidth, and typical workload. Compare dense FP16 against memory bandwidth below. AMD and Intel Gaudi are coming soon.

Get Started

Compare / Dense TFLOPS

Bandwidth × compute

Dense figures only — no sparsity doubling. Most serving work is bandwidth-bound in decode and compute-bound in prefill; treat peak FLOPS across vendors as directional, not interchangeable.

Dense FP16 vs memory bandwidth

Bubble = VRAM

  • NVIDIA
  • AMD
  • Intel

Ranked metric

Lead: B300 at 3,500 TFLOPS

NVIDIA

NVIDIA A10 — Ampere Data Center GPU

24 GB

Architecture
Ampere
Generation
Ampere
Bandwidth
600 GB/s
Workload
Inference · Graphics
FP16 dense
125 TFLOPS
Get Started

NVIDIA A10G — Ampere Graphics & Inference GPU

22 GB

Architecture
Ampere
Generation
Ampere
Bandwidth
600 GB/s
Workload
Inference · Graphics
FP16 dense
70 TFLOPS
Get Started

NVIDIA A16 — Ampere Virtualization GPU

16 GB

Also listed: 2 GB / 4 GB / 8 GB / 16 GB / 32 GB / 64 GB / 128 GB

Architecture
Ampere
Generation
Ampere
Bandwidth
200 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA A30 — Ampere Tensor Core GPU

24 GB

Architecture
Ampere
Generation
Ampere
Bandwidth
933 GB/s
Workload
Training · Inference
FP16 dense
165 TFLOPS
Get Started

NVIDIA A40 — Ampere Data Center Graphics GPU

48 GB

Also listed: 2 GB / 4 GB / 8 GB / 12 GB / 16 GB / 24 GB / 48 GB

Architecture
Ampere
Generation
Ampere
Bandwidth
696 GB/s
Workload
Inference · Graphics
FP16 dense
150 TFLOPS
Get Started

NVIDIA A100 — Ampere Tensor Core AI GPU

80 GB

Also listed: 40 GB / 80 GB

Architecture
Ampere
Generation
Ampere
Bandwidth
2.0 TB/s
Workload
Training · Inference
FP16 dense
312 TFLOPS
Get Started

NVIDIA B200 — Blackwell Tensor Core AI GPU

192 GB

Also listed: 179 GB / 180 GB / 192 GB

Architecture
Blackwell
Generation
Blackwell
Bandwidth
8 TB/s
Workload
Training · Inference
FP16 dense
2,250 TFLOPS
Get Started

NVIDIA B300 — Blackwell Ultra Tensor Core AI GPU

288 GB

Also listed: 268 GB / 270 GB / 288 GB

Architecture
Blackwell
Generation
Blackwell
Bandwidth
8 TB/s
Workload
Training · Inference
FP16 dense
3,500 TFLOPS
Get Started

NVIDIA GB200 — Grace Blackwell AI Superchip

192 GB

Architecture
Blackwell
Generation
Grace Blackwell
Bandwidth
8 TB/s
Workload
Training · Inference
FP16 dense
2,500 TFLOPS
Get Started

NVIDIA GeForce RTX 3070 — Ampere Gaming GPU

8 GB

Architecture
Ampere
Generation
GeForce 30
Bandwidth
448 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 3080 — Ampere Gaming GPU

10 GB

Architecture
Ampere
Generation
GeForce 30
Bandwidth
760 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 3080 Ti — Ampere Gaming GPU

12 GB

Architecture
Ampere
Generation
GeForce 30
Bandwidth
912 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 3090 — Ampere Enthusiast GPU

24 GB

Architecture
Ampere
Generation
GeForce 30
Bandwidth
936 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 3090 Ti — Ampere Enthusiast GPU

24 GB

Architecture
Ampere
Generation
GeForce 30
Bandwidth
1.0 TB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 4070 Ti — Ada Lovelace Gaming GPU

12 GB

Architecture
Ada Lovelace
Generation
GeForce 40
Bandwidth
504 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 4080 — Ada Lovelace Gaming GPU

16 GB

Architecture
Ada Lovelace
Generation
GeForce 40
Bandwidth
717 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 4080 SUPER — Ada Lovelace Gaming GPU

16 GB

Architecture
Ada Lovelace
Generation
GeForce 40
Bandwidth
736 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 4090 — Ada Lovelace Enthusiast GPU

24 GB

Architecture
Ada Lovelace
Generation
GeForce 40
Bandwidth
1.0 TB/s
Workload
Graphics
FP16 dense
165 TFLOPS
Get Started

NVIDIA GeForce RTX 5080 — Blackwell Gaming GPU

16 GB

Architecture
Blackwell
Generation
GeForce 50
Bandwidth
960 GB/s
Workload
Graphics
Get Started

NVIDIA GeForce RTX 5090 — Blackwell Enthusiast GPU

32 GB

Architecture
Blackwell
Generation
GeForce 50
Bandwidth
1.8 TB/s
Workload
Graphics
FP16 dense
419 TFLOPS
Get Started

NVIDIA GH200 — Grace Hopper AI Superchip

96 GB

Architecture
Hopper
Generation
Hopper
Bandwidth
4 TB/s
Workload
Training · Inference
FP16 dense
990 TFLOPS
Get Started

NVIDIA H100 — Hopper Tensor Core AI GPU

80 GB

Also listed: 80 GB / 94 GB

Architecture
Hopper
Generation
Hopper
Bandwidth
3.4 TB/s
Workload
Training · Inference
FP16 dense
989 TFLOPS
Get Started

NVIDIA H200 — Hopper Tensor Core AI GPU

141 GB

Also listed: 141 GB / 143 GB

Architecture
Hopper
Generation
Hopper
Bandwidth
4.8 TB/s
Workload
Training · Inference
FP16 dense
989 TFLOPS
Get Started

NVIDIA L4 — Ada Lovelace Inference GPU

24 GB

Also listed: 22 GB / 24 GB

Architecture
Ada Lovelace
Generation
Ada Lovelace
Bandwidth
300 GB/s
Workload
Inference · Graphics
FP16 dense
121 TFLOPS
Get Started

NVIDIA L40 — Ada Lovelace Data Center GPU

48 GB

Architecture
Ada Lovelace
Generation
Ada Lovelace
Bandwidth
864 GB/s
Workload
Inference · Graphics
FP16 dense
181 TFLOPS
Get Started

NVIDIA L40S — Ada Lovelace AI & Graphics GPU

48 GB

Also listed: 44 GB / 48 GB / 96 GB / 192 GB

Architecture
Ada Lovelace
Generation
Ada Lovelace
Bandwidth
864 GB/s
Workload
Inference · Graphics
FP16 dense
181 TFLOPS
Get Started

NVIDIA RTX 2000 Ada — Ada Lovelace Professional GPU

16 GB

Architecture
Ada Lovelace
Generation
RTX Ada
Bandwidth
224 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX 4000 Ada — Ada Lovelace Professional GPU

20 GB

Architecture
Ada Lovelace
Generation
RTX Ada
Bandwidth
360 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX 5000 Ada — Ada Lovelace Professional GPU

32 GB

Architecture
Ada Lovelace
Generation
RTX Ada
Bandwidth
576 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX 6000 Ada — Ada Lovelace Professional GPU

48 GB

Architecture
Ada Lovelace
Generation
RTX Ada
Bandwidth
960 GB/s
Workload
Inference · Graphics
FP16 dense
182 TFLOPS
Get Started

NVIDIA RTX 5000 — Turing Professional GPU

32 GB

Architecture
Turing
Generation
Quadro RTX
Bandwidth
448 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX 6000 — Turing Professional GPU

24 GB

Architecture
Turing
Generation
Quadro RTX
Bandwidth
672 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX A2000 — Ampere Professional GPU

6 GB

Architecture
Ampere
Generation
RTX Ampere
Bandwidth
288 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX A4000 — Ampere Professional GPU

16 GB

Architecture
Ampere
Generation
RTX Ampere
Bandwidth
448 GB/s
Workload
Inference · Graphics
FP16 dense
77 TFLOPS
Get Started

NVIDIA RTX A4500 — Ampere Professional GPU

20 GB

Architecture
Ampere
Generation
RTX Ampere
Bandwidth
640 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX A5000 — Ampere Professional GPU

24 GB

Architecture
Ampere
Generation
RTX Ampere
Bandwidth
768 GB/s
Workload
Inference · Graphics
FP16 dense
111 TFLOPS
Get Started

NVIDIA RTX A6000 — Ampere Professional GPU

48 GB

Architecture
Ampere
Generation
RTX Ampere
Bandwidth
768 GB/s
Workload
Inference · Graphics
FP16 dense
155 TFLOPS
Get Started

NVIDIA RTX PRO 4000 — Blackwell Professional GPU

24 GB

Architecture
Blackwell
Generation
RTX PRO Blackwell
Bandwidth
672 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX PRO 4500 — Blackwell Professional GPU

32 GB

Architecture
Blackwell
Generation
RTX PRO Blackwell
Bandwidth
896 GB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX PRO 5000 — Blackwell Professional GPU

48 GB

Architecture
Blackwell
Generation
RTX PRO Blackwell
Bandwidth
1.3 TB/s
Workload
Inference · Graphics
Get Started

NVIDIA RTX PRO 6000 — Blackwell Professional GPU

96 GB

Architecture
Blackwell
Generation
RTX PRO Blackwell
Bandwidth
1.8 TB/s
Workload
Inference · Graphics
FP16 dense
500 TFLOPS
Get Started

NVIDIA T4 — Turing Tensor Core Inference GPU

16 GB

Architecture
Turing
Generation
Turing
Bandwidth
320 GB/s
Workload
Inference
FP16 dense
65 TFLOPS
Get Started

NVIDIA V100 — Volta Tensor Core AI GPU

32 GB

Also listed: 16 GB / 32 GB

Architecture
Volta
Generation
Volta
Bandwidth
900 GB/s
Workload
Training · Inference
FP16 dense
125 TFLOPS
Get Started

AMD

AMD Instinct MI25 — GCN AI & HPC GPU

Coming Soon

16 GB

Architecture
GCN
Generation
Instinct
Bandwidth
484 GB/s
Workload
Training · Inference
Get Started

AMD Instinct MI300X — CDNA 3 AI Accelerator

Coming Soon

192 GB

Architecture
CDNA 3
Generation
Instinct
Bandwidth
5.3 TB/s
Workload
Training · Inference
FP16 dense
1,307 TFLOPS
Get Started

AMD Instinct MI325X — CDNA 3 AI Accelerator

Coming Soon

256 GB

Architecture
CDNA 3
Generation
Instinct
Bandwidth
6 TB/s
Workload
Training · Inference
FP16 dense
1,307 TFLOPS
Get Started

AMD Instinct MI350X — CDNA 4 AI Accelerator

Coming Soon

288 GB

Architecture
CDNA 4
Generation
Instinct
Bandwidth
8 TB/s
Workload
Training · Inference
FP16 dense
2,300 TFLOPS
Get Started

AMD Instinct MI355X — CDNA 4 AI Accelerator

Coming Soon

288 GB

Architecture
CDNA 4
Generation
Instinct
Bandwidth
8 TB/s
Workload
Training · Inference
FP16 dense
2,500 TFLOPS
Get Started

AMD Radeon Pro V520 — RDNA Professional GPU

Coming Soon

8 GB

Architecture
RDNA
Generation
Radeon Pro
Bandwidth
448 GB/s
Workload
Graphics
Get Started

AMD Radeon Pro V710 — RDNA 3 Professional GPU

Coming Soon

24 GB

Architecture
RDNA 3
Generation
Radeon Pro
Bandwidth
432 GB/s
Workload
Graphics
Get Started

AMD Radeon RX 6700 XT — RDNA 2 Gaming GPU

Coming Soon

12 GB

Architecture
RDNA 2
Generation
Radeon RX 6000
Bandwidth
384 GB/s
Workload
Graphics
Get Started

Intel

Intel Gaudi 2 — Gaudi 2 AI Accelerator

Coming Soon

96 GB

Architecture
Gaudi
Generation
Gaudi 2
Bandwidth
2.5 TB/s
Workload
Training · Inference
FP16 dense
838 TFLOPS
Get Started

Intel Gaudi HL-205 — Gaudi AI Accelerator

Coming Soon

32 GB

Architecture
Gaudi
Generation
Gaudi
Bandwidth
1 TB/s
Workload
Training · Inference
Get Started