NEWS
AMD’s Halo Station Tries to Pull AI Off the Cloud
AMD’s Threadripper Halo Station packs up to 2.6TB of memory for local trillion-parameter models, a 2027 bid to stop cloud inference bills.
AMD unveiled the Threadripper Halo Station at IFA on September 4, 2026, a liquid-cooled tower it says can run AI models with more than a trillion parameters. Jack Huynh, senior vice president and general manager of Computing and Graphics, called it the most powerful workstation in the world. The IFA unit used two Instinct MI350P cards; the headline 576GB of HBM3E needs four.
The company is selling a way off the cloud inference meter for teams that want large models and long-running agents on hardware they own, a category Nvidia already sells, while AMD’s version arrives in 2027 with more memory and no price.
Two Cards on Stage, Four on the Spec Sheet
Huynh never named the processor beyond 96 cores of Threadripper Pro. That flagship is the Ryzen Threadripper PRO 9995WX, a Zen 5 chip with 192 threads, a 2.5 GHz base clock, a 5.4 GHz boost, 384 MB of L3 cache, and a 350W TDP. It talks to eight-channel DDR5-6400 and 128 PCIe 5.0 lanes, and it launched at $11,699 on July 23, 2025.
Beside it sat two liquid-cooled Instinct MI350P cards, each with 144GB of HBM3E at 4 TB/s. Huynh said there is a path to four, which is how the 576GB figure gets on stage. The two-card IFA machine holds 288GB of accelerator memory, half of that ceiling.
THE IFA UNIT VERSUS THE MAX SPEC
- On stage: Two MI350P cards and 288GB of HBM3E in a tower with separate closed-loop radiators on the CPU and each card.
- On the spec sheet: Up to four cards, 576GB of HBM3E, and 2TB of RDIMM system memory.
- The demo: Huynh generated a 3D world and a flight simulator from a single prompt, then said the box should stay quiet under liquid cooling.
The showcase unit sat on an ASUS WRX90 workstation board with an SFP network card fitted. That is a Channel parts list, not a soldered superchip. AMD already has a smaller Halo box on Ryzen AI Max silicon; this tower is the other end of that family, built so Instinct cards can leave the rack.
2.6TB of Split Memory Versus Nvidia’s 748GB Pool
AMD lists up to 2.6TB of combined memory and rates that total at 3.4x the Nvidia DGX Station, with up to 16.4TB/s of system memory bandwidth. The 2.6TB figure is 576GB of HBM3E plus 2TB of RDIMM. Nvidia’s shipping station uses one GB300 Grace Blackwell Ultra superchip, with 252GB of HBM3e at 7.1 TB/s and 496GB of LPDDR5X, for a datasheet total of 748 GB of coherent memory tied together at 900 GB/s over NVLink-C2C.
DESKSIDE MEMORY AND LINKS COMPARED
| Attribute | Halo Station, four-card spec | Nvidia DGX Station |
|---|---|---|
| CPU | 96-core 9995WX, 350W | 72-core Grace |
| HBM | 576GB HBM3E (4 x 144GB) | 252GB HBM3e |
| CPU memory | 2TB DDR5 RDIMM, 410 GB/s | 496GB LPDDR5X, 396 GB/s |
| GPU memory bandwidth | 16 TB/s | 7.1 TB/s |
| CPU-GPU link | PCIe 5.0 x16 per card | NVLink-C2C at 900 GB/s |
A trillion parameters at 4 bits is about 500GB of weights, because 1,000,000,000,000 times 4 bits divided by 8 is 500GB. Four cards hold 576GB of HBM3E, which leaves about 76GB for runtime state if the weights sit on the accelerators. Two cards at 288GB cannot hold that 500GB file. Nvidia also says the DGX Station can run models up to 1 trillion parameters, with quantization, inside a single coherent pool rather than two kinds of memory that the software has to stitch together.
The memory AMD is advertising is larger. It is also split. HBM lives on the cards and RDIMM lives on the CPU, and the MI350P has no GPU scale-up fabric of the kind Nvidia uses inside a superchip. Capacity is the Halo Station’s argument. Coherence is the DGX Station’s.
Why AMD Wants Agents Off the Cloud Meter
AMD wrote the product for people who already pay a cloud bill all day: researchers, model developers, simulation users, and teams that cannot put sensitive data on a shared cluster. The page tells them to run agents as long as they need and keep the money they would spend on cloud inference, which is the job this tower is actually being asked to do.
This is the most powerful workstation in the world, designed and engineered for a complete new era of computing, capable of running AI models with more than a trillion parameters.
Jack Huynh, senior vice president and general manager of Computing and Graphics, IFA keynote
He also told the hall this is about as close as you can get to a personal supercomputer, and that the new workstation will redefine personal AI. Personal, in this class, means the box parks next to a desk rather than in a rented rack. Huynh said the machine is for a complete new era of computing; the customer AMD named is the one already constrained by cloud access or shared infrastructure, including shops that need to keep IP on premises.
Frontier models and fleets of agents change the math because they do not stop when the prompt returns. A rented GPU that runs all afternoon becomes a line item. A tower you own does not zero that cost, but it does cap it, which is why both AMD and Nvidia are suddenly in the deskside business with the same trillion-parameter claim.
Up to 600 Watts per Instinct Card
The MI350P is a CDNA 4 card AMD launched for air-cooled servers, a half-size relative of the MI350X with 128 compute units and a 10.5-inch dual-slot board. Peak MXFP4 throughput is 4.6 PFLOPs. Official board power is up to 600 watts of board power, with a 450W mode for chassis that cannot feed the full envelope. Server cards use a passive heatsink and forced air. A deskside tower does not, so the IFA unit put liquid blocks on cards that were designed to live in a rack.
WHAT EACH MI350P BRINGS
- Memory: 144GB of HBM3E on a 4096-bit bus at 4 TB/s, with full-chip ECC.
- Power: 600W TBP maximum, 450W configurable, fed by a 12V-2×6 connector.
- Link: PCIe 5.0 x16 only; the card does not carry a GPU-to-GPU scale-up interconnect.
- Software: ROCm, HIP, PyTorch, TensorFlow, JAX, Triton, and SGLang on Linux x86_64.
Two 600W cards plus a 350W CPU already total 1,550 watts before memory, storage, and pumps. Four cards raise that silicon load to 2,750 watts. Nvidia rates the entire DGX Station at 1,600 watts. Huynh said the Halo Station is all liquid cooled because the company wants that super power to stay super quiet. Quiet is a relative claim next to a 2,750-watt parts list.
Nvidia’s DGX Station Already Ships
Nvidia’s box is not a prototype. OEM and dealer quotes run from about $85,000 to $126,000, and the datasheet rates it at up to 20 petaFLOPS of AI compute with Ubuntu and CUDA-X already loaded. MIG can split the GPU into as many as 7 isolated instances, so a single tower can serve a small team. Two stations can be linked over ConnectX-8 if one is not enough.
The software pitch is the part AMD cannot photocopy. Code written on a DGX Station is meant to move onto a DGX rack without a rewrite, which is the path shops already know. AMD is asking those same shops to wait until 2027, then run the same class of job on ROCm across four PCIe cards whose memories are not one pool. Local-AI builders who watched the IFA spec sheet fixated on 576GB of HBM because that is the number that makes a trillion-parameter model fit. The quieter problem is getting from that memory map onto a toolchain their clusters already speak.
A 2027 Prototype With No Partners and No Price
AMD’s product page calls the Halo Station a prototype shown for the first time at IFA 2026 and puts shipment in 2027. It names no OEM, no street price, and no operating system for the tower. Independent tests of the trillion-parameter claim have not been published.
WHAT WE KNOW
- The hardware path: A 9995WX plus two MI350P cards on stage, with a stated path to four cards and 576GB of HBM3E.
- The window: Coming in 2027, with liquid cooling on the CPU and the accelerators.
- The customer AMD named: Individuals and small teams now stuck on cloud access or shared infrastructure, plus groups that must keep data on site.
WHAT IS UNCONFIRMED
- Price: AMD has not published one; the 96-core chip alone launched at $11,699.
- Who builds it: No workstation partner has been named, even though the IFA unit used a standard WRX90 board.
- The four-card SKU: It is still unclear whether AMD will sell that setup or only leave slots for it.
- The stack on the box: OS, drivers, and a supported model zoo for this chassis have not been disclosed.
Because the IFA machine was assembled from parts already in the channel, AMD could in theory validate a SKU faster than a custom module would allow. It has still given buyers a year, no invoice, and no badge on the badge plate. Until those three things exist, the cloud-exit desk you can actually purchase is Nvidia’s.
ROCm Must Match Nvidia’s Shipping Software Stack
The MI350P page lists PyTorch, TensorFlow, JAX, Triton, HIP, and SGLang, which is the open stack AMD has been pushing through ROCm. That list is necessary. It is not the same as a preloaded deskside image that already matches the cluster down the hall. Nvidia ships Ubuntu with CUDA-X, NIM microservices, and a documented hop from the desk to the data center. AMD has not said what image boots on Halo Station, and the accelerators themselves are Linux-only on the spec sheet.
Memory will get AMD the meeting. Four cards can hold a 4-bit trillion-parameter weight file that two cards cannot, and 2TB of RDIMM gives the CPU a place to park agent state that would otherwise sit on a cloud bill. The meeting still ends on software, power, and a date. A 2,750-watt four-card build is a different electrical job from a 1,600-watt DGX Station, and PCIe is a different fabric from NVLink-C2C.
Huynh’s IFA machine made a flight simulator from a prompt and put a 96-core Threadripper next to Instinct cards that used to live only in servers. The product page says that setup arrives in 2027. Until then, the deskside tower that already exists, already has a price range, and already speaks CUDA is the one taking the cloud-exit orders.
-
NEWS3 weeks agoRune Bets on Hillerød for His Achilles Comeback
-
NEWS2 weeks agoPermafrost Thaw Runs Fastest in Mountains, Not Arctic Soils
-
NEWS2 weeks agoThe Roman Space Telescope Flies After Four Budget Fights
-
BUSINESS2 weeks agoCargill Restarts Fort Morgan After an 89-Day Lockout
-
BUSINESS2 weeks agoInfluencer Investors Trade Cash Fees for Illiquid Equity
-
BUSINESS1 week agoSoftware Stocks Rally as Underweight Funds Face Dreamforce
-
NEWS4 weeks agoHamilton’s Ferrari Upgrade Pushes Mercedes Into a Monza Bind
-
NEWS2 weeks agoMurphy Defends a British Open Title Nobody Keeps
