Skip to product information
1 of 1

Nvidia

Nvidia Tesla A100 40GB HBM2 5120-Bit PCI-Express 4.0 x16 1x 8-Pin Graphics Card - 699-21001-0200-400

Nvidia Tesla A100 40GB HBM2 5120-Bit PCI-Express 4.0 x16 1x 8-Pin Graphics Card - 699-21001-0200-400

Regular price $5,175.00 USD
Regular price Sale price $5,175.00 USD
Sale Sold out
Shipping calculated at checkout.
Quantity

NVIDIA Tesla A100 40GB HBM2 PCIe 4.0 x16 GPU — 699-21001-0200-400

The NVIDIA Tesla A100 40GB HBM2 (699-21001-0200-400) is an alternate SKU variant of NVIDIA's Ampere-architecture A100 PCIe data center GPU accelerator. The 699-21001-0200-400 part number typically denotes an OEM or channel-specific configuration — verify compatibility with your server platform before ordering. Featuring 40GB of HBM2 memory, 1,555 GB/s memory bandwidth, and a PCIe 4.0 x16 interface, it delivers the same transformative AI training, deep learning inference, and HPC performance as the reference A100 PCIe.

Key Specifications

  • Part Number / MPN: 699-21001-0200-400
  • Manufacturer: NVIDIA
  • Product Line: Tesla A100
  • Architecture: Ampere (GA100)
  • Memory: 40GB HBM2
  • Memory Bandwidth: 1,555 GB/s
  • Memory Bus Width: 5120-bit
  • Interface: PCI-Express 4.0 x16
  • Power Connector: 1x 8-pin PCIe
  • TDP: 250W
  • Cooling: Passive (requires server airflow)
  • Form Factor: Full-height, full-length (FHFL) PCIe
  • SKU Type: OEM/channel variant — verify server compatibility

Compute Performance

  • FP64 (Double Precision): 9.7 TFLOPS
  • FP64 Tensor Core: 19.5 TFLOPS
  • FP32 (Single Precision): 19.5 TFLOPS
  • TF32 Tensor Core: 156 TFLOPS (312 TFLOPS with sparsity)
  • BFLOAT16 Tensor Core: 312 TFLOPS (624 TFLOPS with sparsity)
  • FP16 Tensor Core: 312 TFLOPS (624 TFLOPS with sparsity)
  • INT8 Tensor Core: 624 TOPS (1,248 TOPS with sparsity)
  • Memory Bandwidth: 1,555 GB/s

Memory

  • Capacity: 40GB HBM2
  • Bandwidth: 1,555 GB/s
  • Bus Width: 5120-bit
  • ECC: Yes — hardware ECC for data integrity

Features

  • NVIDIA Ampere architecture with 3rd-gen Tensor Cores
  • Structural sparsity support — up to 2x throughput on sparse AI models
  • Multi-Instance GPU (MIG) — partition into up to 7 isolated GPU instances
  • NVLink 3.0 support (via NVLink Bridge for multi-GPU PCIe configurations)
  • PCIe 4.0 x16 — compatible with PCIe 3.0 servers at reduced bandwidth
  • Passive cooling — designed for high-airflow data center environments
  • Supports CUDA 11+, cuDNN, TensorRT, NCCL, and NVIDIA AI Enterprise

Physical & Environmental

  • Form Factor: Full-height, full-length (FHFL) dual-slot PCIe
  • TDP: 250W
  • Power Connector: 1x 8-pin PCIe
  • Cooling: Passive (server airflow required)
  • Operating Temperature: 0°C to 35°C (inlet air)

Compatibility

  • Compatible with PCIe 4.0 x16 and PCIe 3.0 x16 server slots
  • OEM/channel variant — verify specific server platform qualification before ordering
  • Requires 250W slot/system power budget per GPU
  • Compatible with NVIDIA AI Enterprise, CUDA 11.x/12.x, PyTorch, TensorFlow, JAX, MXNet
  • Ideal for AI/ML training, LLM inference, HPC simulation, genomics, financial modeling, and scientific computing
View full details