If you are building or scaling an AI system in 2026, chances are you have already encountered the NVIDIA H100 GPU. It is the chip powering the majority of large language model training runs, production inference pipelines for 70B+ parameter models, and enterprise AI infrastructure worldwide. Whether you are researching H100 cloud GPU options, comparing H100 SXM5 vs PCIe configurations, or looking for H100 pricing in India, this guide covers all of it.
Below you will find the complete H100 spec breakdown, real-world performance context, a clear comparison of the two variants, workload-based recommendations, and how Indian AI teams can access H100 GPUs on demand through Race Engineering with INR billing, no dollar conversion, and sub-2-minute provisioning.
What Is the NVIDIA H100 GPU?
The NVIDIA H100 is a data-center GPU built on the Hopper architecture using the GH100 die, manufactured on TSMC's 4N process. Released in 2022, it succeeded the A100 and delivered a generational leap in AI compute performance, particularly for transformer-based models like GPT, LLaMA, Gemini, and Mistral.
Three features define the H100 above all others. First, 80 GB of HBM memory enough to hold large models entirely in VRAM at full precision. Second, the Transformer Engine with FP8 support, which dynamically switches precision mid-computation to maximise throughput without sacrificing accuracy. Third, NVLink 4.0 on the SXM variant, delivering 900 GB/s bidirectional bandwidth between GPUs and enabling efficient multi-GPU training clusters.
Together, these make the H100 the de facto standard for anyone training or serving models above 30 billion parameters.
NVIDIA H100 Full Specs: SXM5 vs PCIe
The H100 comes in two variants with meaningfully different performance profiles. Both are built on the Hopper GH100 architecture, but they differ significantly in memory, bandwidth, power draw, and connectivity.
The H100 SXM5 carries 80 GB of HBM3 memory with 3.35 TB/s of memory bandwidth. It delivers 3,958 TFLOPS of FP8 throughput with sparsity and 989 TFLOPS of TF32 throughput with sparsity. Its TDP is 700W, and it includes NVLink 4.0 at 900 GB/s bidirectional bandwidth, with MIG support for up to 7 isolated instances. On Race Engineering, the H100 SXM5 is available at ₹290/hr with 80 GB HBM3, 200 GB RAM, and 16 vCPUs.
The H100 PCIe carries 80 GB of HBM2e memory with approximately 2.0 TB/s of memory bandwidth. It delivers 3,026 TFLOPS of FP8 throughput with sparsity and 756 TFLOPS of TF32 throughput with sparsity. Its TDP is 350W, half that of the SXM5, and it does not include NVLink support. It also supports MIG partitioning into up to 7 instances and is available on Race Engineering on request.
H100 SXM5 vs PCIe: Which One Do You Need?
Choose the SXM5 for large-scale training and 70B+ inference. The SXM5 variant is the high-performance option. Its 3.35 TB/s HBM3 memory bandwidth is 68 percent faster than the PCIe variant, and NVLink 4.0 support allows you to build multi-GPU clusters where cards communicate at 900 GB/s, far beyond what PCIe interconnects can offer. For distributed training of foundation models, multi-node LLM serving, or any workload requiring tight GPU-to-GPU communication, SXM5 is the correct choice.
Choose the PCIe variant for inference and single-GPU fine-tuning. It fits into standard server slots at 350W, half the power draw of the SXM5, and skips the need for specialised DGX or HGX server platforms. It still carries 80 GB of HBM2e memory, supports MIG partitioning into up to 7 isolated instances, and delivers strong inference performance. For teams running single-GPU fine-tuning jobs or serving mid-size models without multi-GPU requirements, the PCIe variant offers a more accessible and cost-effective path.
The rule of thumb is simple: if your workload needs multi-GPU NVLink scaling or trains models above 30B parameters, choose SXM5. For everything else, PCIe delivers the same 80 GB VRAM at a lower cost and infrastructure complexity.
H100 Performance: What It Actually Delivers
For LLM training, the H100's Transformer Engine is the headline feature. By automatically switching between FP8 and FP16 precision layer by layer, it delivers roughly 3 to 4 times the training throughput of the A100 on transformer-based models. Real-world training runs for 70B parameter models have completed 2 to 3 times faster on H100 clusters compared to equivalent A100 setups. For teams with long training schedules, that compression in compute time is the H100's strongest argument.
For LLM inference, the 3.35 TB/s memory bandwidth of the SXM5 makes it the fastest GPU available for large-model inference. At 70B parameters in FP16, the H100 can serve multiple concurrent requests with sub-second latency, something no consumer GPU or lower-tier data-center card can match at full precision. Multi-Instance GPU support also allows the card to serve multiple isolated tenants simultaneously, improving utilisation economics for inference platforms.
On the question of H100 vs A100, the A100 remains a capable GPU for inference workloads that fit within its capabilities. However, the H100 outperforms it by roughly 3 to 4 times on transformer training via FP8, offers 68 percent more memory bandwidth on the SXM variant, and supports the faster NVLink 4.0 interconnect. For teams actively training large models, the upgrade from A100 to H100 is one of the clearest performance-per-dollar improvements available. At Race Engineering, the A100 SXM4 is available at ₹182/hr, a strong option for teams where H100 throughput is not yet required.
Which Workloads Actually Need an H100?
The H100 is not always the right answer. Matching your workload to the correct GPU tier before spending budget is essential.
For LLM training at 70B parameters and above, the H100 SXM5 in a multi-GPU configuration is the right choice. NVLink, combined with 3.35 TB/s bandwidth, enables fast model parallelism that smaller cards simply cannot deliver.
For LLM inference on models between 30B and 70B parameters, a single H100 SXM5 is ideal. The 80 GB of VRAM and low-latency serving at full FP16 precision make it the standard choice for production inference pipelines.
For fine-tuning models between 7B and 30B parameters, the H100 PCIe or an A100 is the more cost-effective option. Full precision training with LoRA or QLoRA fits comfortably within 80 GB without needing the SXM5's full power.
For Stable Diffusion and image generation workloads, an RTX 4090 or L40S is the better choice. The H100 is overkill for these tasks; save the budget for inference fleet scaling.
For multi-tenant inference serving, the H100 SXM5 with MIG partitioning is purpose-built. Each card can be split into up to 7 isolated instances, making it highly efficient for platforms serving multiple workloads simultaneously.
Rent an H100 GPU in India, Race Engineering
For most AI teams, buying an H100 outright is not practical. A single PCIe card costs $25,000 to $30,000 on the open market. SXM server systems start at $250,000. Lead times are unpredictable, and the hardware depreciates the moment you buy it.
Race Engineering gives Indian AI developers and teams on-demand access to H100 SXM5 GPUs billed in INR, provisioned in under 2 minutes, with no long-term commitment required.
The H100 SXM5 is available at ₹290/hr with 80 GB HBM3, 200 GB RAM, and 16 vCPUs, the most deployed instance on the platform. Billing is INR-native with no dollar conversion, no FX risk, and GST invoices included. Provisioning takes under 2 minutes, the GPU is live and SSH-ready before your coffee cools. Per-minute billing means you can pause at any time, and billing stops instantly, with the ability to resume from the same state. A 99.9% SLA uptime guarantee is backed by a contract with auto credits if missed. Full root access means you can bring your own Docker image, run any framework, and face no restrictions. New users get ₹500 free credits to start with, no card required.
See live H100 availability and pricing at raceengineering.ai/pricing.
Frequently Asked Questions
Q1. How much VRAM does the H100 have?
80 GB on all variants. The SXM5 uses HBM3 with 3.35 TB/s bandwidth. The PCIe version uses HBM2e with around 2 TB/s. There is no 40 GB H100.
Q2. What is the H100 best used for?
Training and serving large language models with above 30B parameters. It is also the standard choice for distributed multi-GPU training, high-throughput production inference, and multi-tenant workloads via MIG partitioning.
Q3. H100 SXM5 or PCIe, which should I pick?
SXM5 if you need NVLink multi-GPU scaling or are training models with above 30B parameters. PCIe if you are running single-GPU fine-tuning or inference, and want lower power draw and infrastructure cost.
Q4. Is the H100 overkill for models under 13B parameters?
Yes, almost always. An RTX 4090 or L40S will handle sub-13B inference at a fraction of the cost. Save the H100 budget for workloads that genuinely need 80 GB of VRAM or NVLink scaling.
Q5. Can I rent an H100 in India without dollar billing?
Yes. Race Engineering offers H100 SXM5 at ₹290/hr with full INR billing, GST invoices, and no dollar conversion. Provisioning takes under 2 minutes. Start at raceengineering.ai



