OpenAI MacStudio AIAgents UnifiedMemory NvidiaRTXSpark

한국어로 읽기 →

Why OpenAI Is Buying Mac Studios by the Rack Instead of GPUs


OpenAI has spent the past several months acquiring Apple Mac minis and Mac Studios in tens-of-thousands-unit volume. Anthropic is renting the same class of Mac infrastructure in bulk through Amazon Web Services. The labs building frontier AI are stacking consumer desktops in their datacenters, rack by rack, rather than the six-figure Nvidia server GPUs you would expect them to reach for.

The buying has moved the supply chain. Retail delivery estimates for Mac Studios configured with large memory have stretched from roughly two weeks to as long as two months. With concentrated corporate demand straining supply, Apple brought its Mac Studio refresh — including an M6 Mac mini and an M5 Ultra option — ahead of its usual announcement cadence.

Why independent desktops instead of big GPU servers

Pretraining a large language model with hundreds of billions of parameters requires very expensive GPU clusters built for raw floating-point throughput. The workloads OpenAI and Anthropic are now pushing hardest on — computer-use agents and reinforcement-learning pipelines — have a different infrastructure profile.

A computer-use agent perceives an operating-system screen the way a person does and drives real software through mouse and keyboard input. Training one at scale means tens of thousands of independent desktop environments running in parallel at the same time. Set against the cost of spinning up virtualized instances on top of premium server GPUs, building that agent-training farm out of Mac minis and Mac Studios is markedly cheaper both to stand up and to operate.

Unified memory and Thunderbolt 5

What makes the Mac the default box here is Apple silicon’s unified memory architecture and its low-latency interconnect.

A typical x86 PC keeps the CPU and GPU physically separate, with data crossing the PCIe bus and creating a transfer bottleneck. On Apple silicon, the CPU, GPU and Neural Engine share a single high-bandwidth memory pool directly. That fits a lightweight agent model in the billions-to-tens-of-billions parameter range: it can sit entirely in memory while it loops between inference and system control without copy latency.

Thunderbolt 5 changes the math on distributed setups. Direct device-to-device links that bypass the heavy standard TCP/IP network stack produce a very low-latency PC cluster without costly InfiniBand or high-end switches. Dozens to hundreds of Macs can be tied together to orchestrate an agent swarm without the network-switching bill that comes with building server racks.

Nvidia’s counter and a $3,400 sellout

As Apple hardware settled in as the cost-effective option for independent agent workloads, Nvidia and the PC makers moved to counter with high-performance AI-PC platforms. The centerpiece is Nvidia’s next-generation superchip, RTX Spark (N1x), slated for a fall release.

Built on TSMC’s 3nm process, N1x integrates a 20-core Grace CPU and a Blackwell RTX 5070 GPU with 6,144 CUDA cores on a single chipset. It runs NVLink-C2C at roughly 600 GB/s between CPU and GPU, carries up to 128 GB of LPDDR5X unified memory, and delivers 1 petaflops of compute at FP4.

Per Taiwan’s Economic Daily, Asus, MSI, Dell, HP, Lenovo and Microsoft have taken the entire first shipment of N1x. Asus’s RTX Spark laptops, including the ProArt P16 and P14, sold out their preorders despite a price around NT$110,000 (about $3,400). Asus co-CEO Hu Shu-bin said channel reservations ran well ahead of plan and that the company is asking Nvidia for additional supply. The read is that the market treats these machines less as creator laptops and more as high-performance agent nodes that run the full Nvidia stack — CUDA, TensorRT — locally.

Memory chipflation and supply-chain pressure

The labs’ scramble for desktops is spilling into the consumer PC component market. Datacenter demand for high-bandwidth memory (HBM) and server DRAM is soaking up general memory supply, and prebuilt high-performance desktops and workstations are themselves being pulled into AI infrastructure. Component prices are climbing as a result. Asus warned on its earnings call that tight memory supply could cap its PC shipments going forward. For consumers, the weight of hardware cost is shifting from the graphics card to high-capacity system memory.

What this isn’t, and what to watch

It would be a stretch to read OpenAI’s Mac buying as a sign that Apple has become an AI-hardware challenger shaking Nvidia’s hold on the datacenter. The central infrastructure for pretraining the largest models is unchanged. This is closer to big-tech labs finding the cheapest hardware detour for one specific parallel workload: computer-use agent training and reinforcement learning.

What would overturn that reading is how quickly Nvidia’s RTX Spark ecosystem takes hold. If Windows-side agent frameworks, backed by full CUDA acceleration, prove higher throughput per unit than Apple silicon’s unified-memory value proposition, the labs’ reliance on Macs could turn out to be a brief transition. If instead Thunderbolt 5 Mac clustering proves out and memory-bandwidth gains on the M5 Ultra and M6 lines arrive quickly, desktop nodes for agent farms harden into a standalone standard in the AI datacenter.

Two things are worth checking directly: the simulation throughput per unit cost of Nvidia’s RTX Spark desktop (GR1X), due in the fourth quarter, and how far Apple actually compresses its hardware supply cycle.

Sources checked

Disclaimer — This article is for general informational purposes only and is not a recommendation to invest in any specific security or product. Investment decisions and their consequences are solely the reader's responsibility.

Add TrueMyPick as a preferred source on Google.