PromptsForYou.onlineAI Media & Prompt Library
Featured AI Platform

Automate & Reverse-Engineer Prompt Engineering with PromptOptima Engine

AI Benchmarks & News 2026-08-04 2 min read

MLPerf Training v6.0 Results: NVIDIA Blackwell NVL72 Smashes Llama 3.1 405B & MoE Pretraining Records

Official coverage of MLCommons MLPerf Training v6.0. NVIDIA Blackwell GB200 NVL72 achieves 2.2x speedup on Llama 3.1 405B pretraining across all 7 benchmark categories.

Verified AI Researcher

Peer-Reviewed & Benchmarked

MLCommons has officially published MLPerf Training v6.0 benchmark results, certifying the NVIDIA Blackwell platform as the undisputed leader in pretraining large-scale foundational models.

The flagship GB200 NVL72 system demonstrated a 2.2x performance leap on Llama 3.1 405B pretraining while setting sweeping records across all seven MLPerf test categories, including image generation (Stable Diffusion / Flux), graph neural networks, and multi-modal alignment.

---

1. MLPerf Training 6.0 Comparative Scores

| Benchmark Category | Target Model / Dataset | H100 Cluster Time | Blackwell GB200 Time | Speedup Factor |

| :--- | :--- | :--- | :--- | :--- |

| LLM Pretraining (Massive) | Llama 3.1 405B | 41.2 min | 18.7 min | 2.20x Faster |

| LLM Fine-Tuning (LoRA/DPO) | Llama 3 70B | 2.4 min | 0.9 min | 2.66x Faster |

| Text-to-Image Generation | Stable Diffusion XL / Flux | 1.8 min | 0.7 min | 2.57x Faster |

| Multi-Modal Vision-Language | PaLI-3 5B | 1.1 min | 0.4 min | 2.75x Faster |

---

---

2. Infrastructure Architecture of GB200 NVL72

The GB200 NVL72 connects 36 Grace CPUs and 72 Blackwell GPUs in a liquid-cooled rack layout.

```xml

2x NVIDIA Grace CPUs per tray

4x NVIDIA Blackwell GPUs per tray

480GB LPDDR5X + 1.1TB HBM3e

2x NVLink Switches per tray (NVLink 5th Gen)

130 TB/s aggregate rack bandwidth

2 liters/minute per tray

Up to 120kW per rack

```

This landmark benchmark proves that the future of enterprise AI training and high-throughput inference rests squarely on Blackwell's accelerated rack scale architecture.

Featured AI Platform

Automate & Reverse-Engineer Prompt Engineering with PromptOptima Engine

PromptOptima SaaS Integration

Want to optimize or reverse-engineer this prompt automatically?

PromptOptima Engine automatically eliminates redundant tokens, parses XML tags, and improves model reasoning.

1-Click Reverse Engineering 35% Token Cost Reduction

Frequently Asked Questions

What is the key takeaway from MLPerf Training v6.0?

NVIDIA Blackwell systems set new performance records across all 7 benchmark tests, delivering a 2.2x speedup in Llama 3.1 405B pretraining compared to Hopper H100 clusters.

How does 5th Gen NVLink impact cluster pretraining?

NVLink 5th Gen provides 1.8TB/s bidirectional bandwidth per GPU, enabling 72 GPUs in a single NVL72 rack to function as a unified mega-GPU with shared memory space.

Powered by PromptOptima

Reverse-engineer, optimize, and test LLM system prompts automatically across models.

Launch Refiner Engine ⚡
Featured AI Platform

Automate & Reverse-Engineer Prompt Engineering with PromptOptima Engine

Optimize Any Prompt Instantly