AI & Compute Roundup: NVIDIA’s Vera Rubin Posts First MLPerf Numbers Next to AMD’s Biggest Cluster Yet

NVIDIA's first Vera Rubin numbers land in MLPerf Inference v6.1 alongside AMD's biggest GPU cluster submission yet. Also this week: Apple's NVLink talks with NVIDIA, and a laptop with 160GB of usable VRAM.

Vera Rubin’s First MLPerf Numbers Land Next to AMD’s Biggest Cluster Yet

MLCommons published MLPerf Inference v6.1 on September 16, and it’s the first real look at numbers from NVIDIA’s next-generation Vera Rubin NVL72 rack. In its debut submission, Vera Rubin NVL72 hit up to 3.7x the throughput of a GB300 NVL72 rack on the Qwen3-VL benchmark, running vLLM with NVIDIA’s Dynamo inference framework. On DeepSeek-R1, using TensorRT-LLM, the gain was up to 2.5x.

Those numbers are NVIDIA’s own, and they’re preview figures ahead of full production deployment, not a shipping product’s final scores. NVIDIA credits most of the gain to two things: NVFP4 precision cutting the memory footprint, and sixth-generation NVLink cutting the latency between GPUs.

AMD showed up with its largest MLPerf submission to date, a 512-GPU Instinct MI355X cluster that led both the offline and server scenarios on DeepSeek-R1 and GPT-OSS-120B, based on Wccftech’s read of the published results. NVIDIA’s own GB300 NVL72 racks, tested at up to 288 GPUs, landed close behind AMD on DeepSeek-R1, but also managed 99 percent scaling efficiency, meaning almost no throughput lost going from one rack up to four. Intel entered this round too, with Arc Pro B70 workstation cards and Xeon systems, smaller submissions, but enough to put a third vendor’s numbers on the same board this cycle.

Worth remembering: none of these are independent, third-party measurements. Every vendor submits its own hardware and software stack under MLCommons’ standardized rules, and MLCommons audits the methodology rather than declaring a winner. Read that way, this round says less about which chip is fastest and more about how each company is choosing to compete right now: NVIDIA on a brand-new architecture still in preview, AMD on cluster scale with silicon that’s already proven, Intel just getting into the conversation at all.

Apple Reportedly in Talks With NVIDIA Over the Interconnect for Its First Server Since 2011

Apple is reportedly building a rack server around its own M-series Ultra silicon, and reportedly weighing NVIDIA’s NVLink Fusion to connect it, which would mean Apple’s own chips talking to each other over an NVIDIA-designed backplane. MacRumors, citing The Information, reports two configurations on the table, built around two or four future M8 Ultra chips, with a possible release in 2029. Apple hasn’t built a rack server since it discontinued Xserve back in 2011, and the reporting frames this one as genuinely reversible: it could ship without NVLink Fusion, or not ship at all.

Wccftech’s own reporting on the same project ties it to Apple’s in-development “Baltra” AI ASIC, a Broadcom-assisted chip reportedly built on TSMC’s N3E process, and describes NVLink Fusion access as giving Apple’s Private Compute inference architecture more routing options between on-device and cloud processing.

The interest itself tells you something. Apple’s M-series Ultra chips already handle local AI workloads well enough that OpenAI and Anthropic are reportedly renting Mac minis and Mac Studios by the “tens of thousands” for reinforcement-learning training, according to The Information’s reporting as relayed by MacRumors. NVLink Fusion is NVIDIA’s answer to that same pressure, coming from the opposite direction: an interconnect that lets a non-NVIDIA chip join a rack built around NVIDIA’s own switching fabric.

HP’s ZBook Ultra G3a Can Turn 160GB of System Memory Into VRAM

HP’s own announcement details the ZBook Ultra G3a mobile workstation, due on shelves in October, which packs up to 192GB of unified system memory, and up to 160GB of that can be allocated as VRAM through HP’s own software utility or BIOS settings, with no discrete workstation GPU required. The chip doing the work is AMD’s flagship Ryzen AI Max+ Pro 495, a 16-core Zen 5 “Gorgon Halo” part with Radeon 8065S integrated graphics and a 55-TOPS NPU, for up to 131 TOPS combined across the whole package.

HP ZBook Ultra G3a mobile workstation
The HP ZBook Ultra G3a, announced September 15, 2026, can allocate up to 160GB of its unified system memory as VRAM. Image: HP Inc.

That VRAM headroom is the real story for anyone running local models. HP and AMD say the configuration supports models up to roughly 300 billion parameters running locally, well beyond what any single consumer discrete GPU’s VRAM allows today. The tradeoff is bandwidth and thermal budget, not capacity: unified LPDDR5X is slower than dedicated GDDR or HBM, and HP’s redesigned thermal system, which the company says delivers up to 81.8% higher sustained TDP than its predecessor, exists specifically to sustain up to 100W of workload inside a 17.9mm, 4.2-pound chassis. HP has not announced pricing.

Sources