Home

Local AI hardware

Ready AI systems and personal AI supercomputers

These are complete systems, not part-by-part builds. Compare current Apple Silicon Mac mini and Mac Studio systems, NVIDIA GB10 AI boxes, Ryzen AI Max devices and large-memory workstations — with transparent memory-fit and software-support reports for local AI workloads.

Memory terminology matters: unified or coherent memory is shared by the CPU and GPU. It is not the same as a graphics card with the same amount of dedicated VRAM. GPU-addressable and per-GPU limits are shown separately where the manufacturer publishes them.

ASUS · Mobile AI workstation

ProArt P16 H7607 (128 GB)

Current generationAnnounced
Compute platform
NVIDIA RTX Spark N1X Superchip
Memory type
Up to 128 GB LPDDR5X coherent unified memory shared by CPU and GPU
AI engine
5th-generation Tensor Cores
Storage / OS
512 GB, 1 TB, or 2 TB M.2 NVMe PCIe 4.0 SSD; two M.2 PCIe 5.0 x4 slots · Windows 11

AI capability: Up to 1 PFLOP FP4 AI performance (manufacturer claim)

Model/workload scope: ASUS positions the 128 GB configuration for local 120B-parameter LLMs, 4K AI video and 90 GB+ 3D scenes

Best for: Portable local AI · AI video generation · 3D creation · CUDA development

ASUS published the H7607 specification at IFA on September 2, 2026; regional pricing and sale dates remain unconfirmed. The unified memory pool is not dedicated VRAM.

Official product page →

ASUS · Mobile AI workstation

ProArt P14 H7407 (128 GB)

Current generationAnnounced
Compute platform
NVIDIA RTX Spark N1X Superchip
Memory type
Up to 128 GB LPDDR5X coherent unified memory shared by CPU and GPU
AI engine
5th-generation Tensor Cores
Storage / OS
512 GB or 1 TB M.2 NVMe PCIe 4.0 SSD; one M.2 PCIe 4.0 x4 slot · Windows 11

AI capability: Up to 1 PFLOP FP4 AI performance (manufacturer claim)

Model/workload scope: ASUS positions the 128 GB configuration for local 120B-parameter LLMs and demanding creator workloads

Best for: Portable local AI · Creator workflows · Local LLM inference · CUDA development

ASUS published the H7407 specification at IFA on September 2, 2026; regional pricing and sale dates remain unconfirmed. The unified memory pool is not dedicated VRAM.

Official product page →

ASUS · Compact AI desktop

ProArt GR1X Mini PC (128 GB)

Current generationAnnounced
Compute platform
NVIDIA RTX Spark N1X Superchip
Memory type
Up to 128 GB LPDDR5X coherent unified memory shared by CPU and GPU
AI engine
5th-generation Tensor Cores
Storage / OS
M.2 NVMe PCIe 5.0 / 4.0 SSD slots; installed capacity not yet published · Windows 11

AI capability: Up to 1 PFLOP FP4 AI performance (manufacturer claim)

Model/workload scope: ASUS positions the 128 GB system for local 120B-parameter LLMs, personal agents and 90 GB+ 3D scenes

Best for: Always-on local agents · Local LLM inference · AI creation · Compact CUDA workstation

ASUS announced the GR1X at IFA on September 2, 2026 and currently offers a notification signup rather than a confirmed sale date. The unified memory pool is not dedicated VRAM.

Official product page →

Lenovo · Mobile AI workstation

Yoga Pro 9n 15-inch (128 GB)

Current generationAnnounced
Compute platform
NVIDIA RTX Spark N1X Superchip
Memory type
Up to 128 GB LPDDR5X coherent unified memory shared by CPU and GPU
AI engine
5th-generation Tensor Cores
Storage / OS
Two SSD slots, up to 4 TB total: slot 1 supports 512 GB or 1 TB PCIe Gen 4 M.2 2242; slot 2 supports 2 TB PCIe Gen 5 M.2 2242 · Windows 11

AI capability: Up to 1 PFLOP FP4 AI performance (manufacturer claim)

Model/workload scope: Large local AI and creator workloads within 128 GB unified-memory limits; exact model fit depends on runtime and quantization

Best for: Portable AI development · Creator workflows · RTX gaming · Local agents

Lenovo announced the Yoga Pro 9n on September 3, 2026 and states that pricing and availability will be announced later. The unified memory pool is not dedicated VRAM.

Official product page →

Acer · Compact AI desktop

Veriton RI110 AI Mini Workstation (96 GB)

Current generationAnnounced
Compute platform
Intel Core Ultra X7 358H
Memory type
Up to 96 GB dual-channel LPDDR5X system memory shared with the integrated Arc B390 GPU
AI engine
Integrated Intel NPU and Arc AI engines
Storage / OS
Up to 4 TB M.2 2280 PCIe Gen 4 SSD · Windows 11 Pro, Windows 11 Home

AI capability: Acer claims support for local AI models up to 120B parameters; no normalized whole-system TOPS figure is published

Model/workload scope: Acer positions the system for local design, inference, content creation and models up to 120B parameters

Best for: Hybrid local AI · Edge AI development · Content creation · Compact office workstation

Acer announced the RI110 on September 2, 2026 for North America Q4 2026 and EMEA Q1 2027. Its 96 GB LPDDR5X is shared system memory, not dedicated VRAM.

Official product page →

Lenovo · Compact AI desktop

ThinkCentre X Ultra (128 GB)

Current generationAnnounced
Compute platform
AMD Ryzen AI Max+ PRO 495
Memory type
Up to 128 GB LPDDR5X-8533 unified system memory; up to 96 GB can be allocated to integrated graphics
AI engine
AMD XDNA NPU, up to 55 TOPS
Storage / OS
Up to 2× 4 TB M.2 2280 SSD; platform slots are PCIe Gen 4 with Gen 5 SSD options listed · Windows 11 Pro, Windows 11 Home, Linux, Ubuntu (certification only), AMD AI OS

AI capability: Up to 55 NPU TOPS; Radeon 8065S integrated GPU for local AI and graphics workloads

Model/workload scope: Lenovo positions one system for substantial local AI and up to four clustered systems for larger models and multi-agent workloads; no exact parameter ceiling is claimed

Best for: Enterprise local AI · Multi-agent workflows · Compact AI development · Clustered inference

Lenovo announced the 1.6 L ThinkCentre X Ultra on September 3, 2026 for November 2026 availability starting at €3,100. The 96 GB graphics allocation comes from 128 GB unified memory and is not dedicated VRAM.

Official product page →

Apple · Compact AI desktop

Mac mini with M6 (32 GB)

Current generationAnnounced
Compute platform
Apple M6
Memory type
Up to 32 GB unified memory shared across CPU, GPU and Neural Engine (up to 170 GB/s)
AI engine
Dual 16-core Neural Engine + GPU Neural Accelerators
Storage / OS
256 GB, 512 GB, 1 TB, or 2 TB SSD · macOS 27

AI capability: Apple claims up to 4× faster AI performance than Mac mini with M4 and up to 13.5× faster LM Studio prompt processing than Mac mini with M1

Model/workload scope: Compact and medium local models where MLX, Core AI/Core ML, llama.cpp or Metal support is available; no exact parameter limit is claimed

Best for: Always-on local agents · Compact local LLMs · Speech and vision AI · Private on-device workflows

Apple announced this configuration on August 25, 2026; pre-orders are open and availability begins September 22. The 32 GB is unified memory, not dedicated VRAM, and Apple publishes no normalized TOPS figure.

AI model compatibility report

Planning guidance only: memory fit and general framework support are separate. The exact model artifact, runtime version, context and workload must be tested before claiming it runs; no speed benchmark is implied.

Conservative usable memory: 24 GB

7–8B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–8B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Context length and KV cache consume memory beyond model weights.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

14B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 14B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Tokenizer, context and quantization format change actual memory use.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

30–35B LLM, 4-bit

Tight memory fitFramework supported

30–35B LLM, 4-bit is a tight memory fit; keep context, batch size and related settings conservative. General framework support exists; the exact model artifact has not been verified on this device.

Long context can move a memory-fit model into swap or failure.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

70B LLM, 4-bit

Insufficient safe memoryFramework supported

There is no safe memory headroom for 70B LLM, 4-bit. General framework support exists; the exact model artifact has not been verified on this device.

This is a memory-fit estimate, not a speed claim. Context and runtime overhead matter.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

100B+ dense or MoE model, 4-bit class

Insufficient safe memoryExperimental support

There is no safe memory headroom for 100B+ dense or MoE model, 4-bit class. Software support is experimental and depends on the model and framework version.

MoE architectures vary widely; verify the exact artifact and active-expert implementation.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · class-v1

SDXL-class image generation

Comfortable memory fitFramework supported

Comfortable memory headroom for SDXL-class image generation; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Resolution, ControlNet, batch size and app implementation change peak memory.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · General framework documentation · sdxl-class-v1

FLUX-class quantized image generation

Tight memory fitExperimental support

FLUX-class quantized image generation is a tight memory fit; keep context, batch size and related settings conservative. Software support is experimental and depends on the model and framework version.

Compatibility depends on the exact model build and framework; CUDA-only nodes remain unavailable.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · flux-class-v1

Whisper large-class transcription

Comfortable memory fitFramework supported

Comfortable memory headroom for Whisper large-class transcription; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Language, batch size and implementation affect throughput.

Frameworks: MLX Whisper · Core ML builds · PyTorch MPS

Evidence: Memory estimate · General framework documentation · whisper-large-class-v1

7–14B LoRA / QLoRA fine-tuning

Insufficient safe memoryFramework supported

There is no safe memory headroom for 7–14B LoRA / QLoRA fine-tuning. General framework support exists; the exact model artifact has not been verified on this device.

Dataset, optimizer states, sequence length and rank can raise memory use substantially.

Frameworks: MLX · PyTorch MPS when operations are supported

Evidence: Memory estimate · General framework documentation · lora-class-v1

Local diffusion video generation

Insufficient safe memoryExperimental support

There is no safe memory headroom for Local diffusion video generation. Software support is experimental and depends on the model and framework version.

Many video pipelines and custom nodes require CUDA; memory fit alone does not prove compatibility.

Frameworks: Model-specific Metal/Core ML/MLX ports

Evidence: Memory estimate · General framework documentation · video-class-v1

CUDA-only models and custom nodes

Comfortable memory fitUnsupported

Comfortable memory headroom for CUDA-only models and custom nodes; speed has not been measured. The required software ecosystem is unsupported on this platform.

Apple Silicon does not support CUDA. Use a Metal/MLX/Core ML alternative or NVIDIA hardware.

Frameworks: NVIDIA CUDA

Evidence: General framework documentation · cuda-ecosystem-v1

Official product page →

Apple · Compact AI desktop

Mac mini with M5 Pro (64 GB)

Current generationAnnounced
Compute platform
Apple M5 Pro
Memory type
Up to 64 GB unified memory shared across CPU, GPU and Neural Engine (307 GB/s)
AI engine
16-core Neural Engine + Neural Accelerator in each GPU core
Storage / OS
512 GB, 1 TB, 2 TB, 4 TB, or 8 TB SSD · macOS 27

AI capability: Apple positions the system for larger local AI models, advanced upscaling and large diffusion workflows

Model/workload scope: Larger quantized local models and diffusion workloads where the exact framework supports Apple silicon

Best for: Local LLM inference · Diffusion image generation · Always-on AI agents · LoRA experimentation

Apple announced this configuration on August 25, 2026; pre-orders are open and availability begins September 22. The 64 GB is unified memory, not dedicated VRAM, and the selected configuration has 307 GB/s bandwidth.

AI model compatibility report

Planning guidance only: memory fit and general framework support are separate. The exact model artifact, runtime version, context and workload must be tested before claiming it runs; no speed benchmark is implied.

Conservative usable memory: 51 GB

7–8B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–8B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Context length and KV cache consume memory beyond model weights.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

14B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 14B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Tokenizer, context and quantization format change actual memory use.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

30–35B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 30–35B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Long context can move a memory-fit model into swap or failure.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

70B LLM, 4-bit

Tight memory fitFramework supported

70B LLM, 4-bit is a tight memory fit; keep context, batch size and related settings conservative. General framework support exists; the exact model artifact has not been verified on this device.

This is a memory-fit estimate, not a speed claim. Context and runtime overhead matter.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

100B+ dense or MoE model, 4-bit class

Insufficient safe memoryExperimental support

There is no safe memory headroom for 100B+ dense or MoE model, 4-bit class. Software support is experimental and depends on the model and framework version.

MoE architectures vary widely; verify the exact artifact and active-expert implementation.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · class-v1

SDXL-class image generation

Comfortable memory fitFramework supported

Comfortable memory headroom for SDXL-class image generation; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Resolution, ControlNet, batch size and app implementation change peak memory.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · General framework documentation · sdxl-class-v1

FLUX-class quantized image generation

Comfortable memory fitExperimental support

Comfortable memory headroom for FLUX-class quantized image generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Compatibility depends on the exact model build and framework; CUDA-only nodes remain unavailable.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · flux-class-v1

Whisper large-class transcription

Comfortable memory fitFramework supported

Comfortable memory headroom for Whisper large-class transcription; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Language, batch size and implementation affect throughput.

Frameworks: MLX Whisper · Core ML builds · PyTorch MPS

Evidence: Memory estimate · General framework documentation · whisper-large-class-v1

7–14B LoRA / QLoRA fine-tuning

Expected to fitFramework supported

7–14B LoRA / QLoRA fine-tuning is expected to fit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Dataset, optimizer states, sequence length and rank can raise memory use substantially.

Frameworks: MLX · PyTorch MPS when operations are supported

Evidence: Memory estimate · General framework documentation · lora-class-v1

Local diffusion video generation

Tight memory fitExperimental support

Local diffusion video generation is a tight memory fit; keep context, batch size and related settings conservative. Software support is experimental and depends on the model and framework version.

Many video pipelines and custom nodes require CUDA; memory fit alone does not prove compatibility.

Frameworks: Model-specific Metal/Core ML/MLX ports

Evidence: Memory estimate · General framework documentation · video-class-v1

CUDA-only models and custom nodes

Comfortable memory fitUnsupported

Comfortable memory headroom for CUDA-only models and custom nodes; speed has not been measured. The required software ecosystem is unsupported on this platform.

Apple Silicon does not support CUDA. Use a Metal/MLX/Core ML alternative or NVIDIA hardware.

Frameworks: NVIDIA CUDA

Evidence: General framework documentation · cuda-ecosystem-v1

Official product page →

Apple · Desktop AI workstation

Mac Studio with M5 Max (128 GB)

Current generationAnnounced
Compute platform
Apple M5 Max
Memory type
Up to 128 GB unified memory shared across CPU, GPU and Neural Engine (up to 614 GB/s)
AI engine
16-core Neural Engine + Neural Accelerator in each GPU core
Storage / OS
512 GB, 1 TB, 2 TB, 4 TB, or 8 TB SSD · macOS 27

AI capability: Apple claims up to 3.9× faster LM Studio prompt processing and 3.5× faster text-to-image performance than M4 Max

Model/workload scope: Large quantized local models, image generation and supported local video workflows without an exact model-size guarantee

Best for: Large local LLM inference · Creative AI · Image generation · Model development and LoRA

Apple announced this configuration on August 25, 2026; pre-orders are open and availability begins September 22. The 128 GB is unified memory, not dedicated VRAM; the 40-core GPU configuration provides 614 GB/s bandwidth.

AI model compatibility report

Planning guidance only: memory fit and general framework support are separate. The exact model artifact, runtime version, context and workload must be tested before claiming it runs; no speed benchmark is implied.

Conservative usable memory: 102 GB

7–8B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–8B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Context length and KV cache consume memory beyond model weights.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

14B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 14B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Tokenizer, context and quantization format change actual memory use.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

30–35B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 30–35B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Long context can move a memory-fit model into swap or failure.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

70B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 70B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

This is a memory-fit estimate, not a speed claim. Context and runtime overhead matter.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

100B+ dense or MoE model, 4-bit class

Expected to fitExperimental support

100B+ dense or MoE model, 4-bit class is expected to fit; speed has not been measured. Software support is experimental and depends on the model and framework version.

MoE architectures vary widely; verify the exact artifact and active-expert implementation.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · class-v1

SDXL-class image generation

Comfortable memory fitFramework supported

Comfortable memory headroom for SDXL-class image generation; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Resolution, ControlNet, batch size and app implementation change peak memory.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · General framework documentation · sdxl-class-v1

FLUX-class quantized image generation

Comfortable memory fitExperimental support

Comfortable memory headroom for FLUX-class quantized image generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Compatibility depends on the exact model build and framework; CUDA-only nodes remain unavailable.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · flux-class-v1

Whisper large-class transcription

Comfortable memory fitFramework supported

Comfortable memory headroom for Whisper large-class transcription; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Language, batch size and implementation affect throughput.

Frameworks: MLX Whisper · Core ML builds · PyTorch MPS

Evidence: Memory estimate · General framework documentation · whisper-large-class-v1

7–14B LoRA / QLoRA fine-tuning

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–14B LoRA / QLoRA fine-tuning; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Dataset, optimizer states, sequence length and rank can raise memory use substantially.

Frameworks: MLX · PyTorch MPS when operations are supported

Evidence: Memory estimate · General framework documentation · lora-class-v1

Local diffusion video generation

Comfortable memory fitExperimental support

Comfortable memory headroom for Local diffusion video generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Many video pipelines and custom nodes require CUDA; memory fit alone does not prove compatibility.

Frameworks: Model-specific Metal/Core ML/MLX ports

Evidence: Memory estimate · General framework documentation · video-class-v1

CUDA-only models and custom nodes

Comfortable memory fitUnsupported

Comfortable memory headroom for CUDA-only models and custom nodes; speed has not been measured. The required software ecosystem is unsupported on this platform.

Apple Silicon does not support CUDA. Use a Metal/MLX/Core ML alternative or NVIDIA hardware.

Frameworks: NVIDIA CUDA

Evidence: General framework documentation · cuda-ecosystem-v1

Official product page →

Apple · Desktop AI workstation

Mac Studio with M5 Ultra (512 GB)

Current generationAnnounced
Compute platform
Apple M5 Ultra
Memory type
Up to 512 GB unified memory shared across CPU, GPU and Neural Engine (1.2 TB/s)
AI engine
32-core Neural Engine + Neural Accelerator in each GPU core
Storage / OS
1 TB, 2 TB, 4 TB, 8 TB, or 16 TB SSD · macOS 27

AI capability: Apple claims up to 4× faster LM Studio prompt processing and 4.3× faster text-to-image performance than M3 Ultra

Model/workload scope: Very large local model and dataset capacity where MLX, Core AI/Core ML or Metal software support exists; no CUDA support

Best for: Very large local inference · Research and fine-tuning · Creative AI · Multi-agent and private AI services

Apple announced this configuration on August 25, 2026. General Mac Studio availability begins September 22, while Apple gives only late October 2026 for the 512 GB configuration. Unified memory is not dedicated VRAM and CUDA is unsupported.

AI model compatibility report

Planning guidance only: memory fit and general framework support are separate. The exact model artifact, runtime version, context and workload must be tested before claiming it runs; no speed benchmark is implied.

Conservative usable memory: 409 GB

7–8B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–8B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Context length and KV cache consume memory beyond model weights.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

14B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 14B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Tokenizer, context and quantization format change actual memory use.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

30–35B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 30–35B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Long context can move a memory-fit model into swap or failure.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

70B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 70B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

This is a memory-fit estimate, not a speed claim. Context and runtime overhead matter.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

100B+ dense or MoE model, 4-bit class

Comfortable memory fitExperimental support

Comfortable memory headroom for 100B+ dense or MoE model, 4-bit class; speed has not been measured. Software support is experimental and depends on the model and framework version.

MoE architectures vary widely; verify the exact artifact and active-expert implementation.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · class-v1

SDXL-class image generation

Comfortable memory fitFramework supported

Comfortable memory headroom for SDXL-class image generation; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Resolution, ControlNet, batch size and app implementation change peak memory.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · General framework documentation · sdxl-class-v1

FLUX-class quantized image generation

Comfortable memory fitExperimental support

Comfortable memory headroom for FLUX-class quantized image generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Compatibility depends on the exact model build and framework; CUDA-only nodes remain unavailable.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · flux-class-v1

Whisper large-class transcription

Comfortable memory fitFramework supported

Comfortable memory headroom for Whisper large-class transcription; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Language, batch size and implementation affect throughput.

Frameworks: MLX Whisper · Core ML builds · PyTorch MPS

Evidence: Memory estimate · General framework documentation · whisper-large-class-v1

7–14B LoRA / QLoRA fine-tuning

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–14B LoRA / QLoRA fine-tuning; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Dataset, optimizer states, sequence length and rank can raise memory use substantially.

Frameworks: MLX · PyTorch MPS when operations are supported

Evidence: Memory estimate · General framework documentation · lora-class-v1

Local diffusion video generation

Comfortable memory fitExperimental support

Comfortable memory headroom for Local diffusion video generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Many video pipelines and custom nodes require CUDA; memory fit alone does not prove compatibility.

Frameworks: Model-specific Metal/Core ML/MLX ports

Evidence: Memory estimate · General framework documentation · video-class-v1

CUDA-only models and custom nodes

Comfortable memory fitUnsupported

Comfortable memory headroom for CUDA-only models and custom nodes; speed has not been measured. The required software ecosystem is unsupported on this platform.

Apple Silicon does not support CUDA. Use a Metal/MLX/Core ML alternative or NVIDIA hardware.

Frameworks: NVIDIA CUDA

Evidence: General framework documentation · cuda-ecosystem-v1

Official product page →

NVIDIA · Edge AI robotics computer

Jetson Orin Nano 2 Developer Kit

Announced
Compute platform
NVIDIA Jetson Orin Nano 2
Memory type
8 GB shared module memory for the integrated Jetson SoC
AI engine
78 TOPS edge AI compute (manufacturer claim)
Storage / OS
Module and developer-kit storage configuration not yet published by NVIDIA · NVIDIA Jetson software stack, Linux

AI capability: 78 TOPS edge AI compute; NVIDIA claims 2× predecessor inference performance

Model/workload scope: Memory-efficient edge inference for compact language, vision-language and robotics models

Best for: Robotics · Vision AI · Drones · Edge inference · Physical AI development

Announced August 25, 2026 for expected first-half 2027 availability. The published 8 GB is shared module memory, not 8 GB dedicated VRAM.

Official product page →

NVIDIA · Personal AI supercomputer

DGX Spark

Available
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x coherent unified system memory (273 GB/s)
AI engine
5th-generation Tensor Cores
Storage / OS
1 TB or 4 TB NVMe M.2 · NVIDIA DGX OS, Linux

AI capability: Up to 1 PFLOP FP4 AI performance (manufacturer claim)

Model/workload scope: Local prototyping, inference and fine-tuning for large models; vendor positions paired systems for larger workloads

Best for: Local LLM inference · Fine-tuning · AI development · CUDA prototyping

Memory is shared by CPU and GPU; it is not 128 GB of dedicated VRAM.

Official product page →

ASUS · Personal AI supercomputer

Ascent GX10

Available
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x unified system memory
AI engine
5th-generation Tensor Cores
Storage / OS
1 TB or 2 TB PCIe 4.0, or 4 TB PCIe 5.0 M.2 NVMe options · NVIDIA DGX OS, Linux

AI capability: Up to 1 PFLOP FP4 AI performance (manufacturer claim)

Model/workload scope: ASUS positions the system for local fine-tuning of models up to 200B parameters

Best for: Local LLM inference · Model fine-tuning · AI agents · Developer workstation

ASUS calls this unified memory; it must not be displayed as dedicated VRAM.

Official product page →

Dell · Personal AI supercomputer

Pro Max with GB10 (FCM1253)

Available
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x coherent unified system memory (273 GB/s)
AI engine
5th-generation Tensor Cores
Storage / OS
1 TB, 2 TB or 4 TB NVMe configurations · NVIDIA DGX OS 7, Linux

AI capability: Up to 1 PFLOP FP4 / 1,000 TOPS sparse AI performance (manufacturer claim)

Model/workload scope: Vendor positions a single system for models up to 200B parameters and paired systems up to 405B

Best for: Local LLM inference · Fine-tuning · AI development · CUDA prototyping

128 GB is one CPU/GPU coherent memory pool, not separate 128 GB RAM plus 128 GB dedicated VRAM.

Official product page →

Lenovo · Personal AI supercomputer

ThinkStation PGX

Available
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x coherent unified system memory (273 GB/s)
AI engine
5th-generation Tensor Cores
Storage / OS
1 TB or 4 TB self-encrypting NVMe options · NVIDIA DGX OS, Ubuntu Linux Pro-based NVIDIA Base OS

AI capability: Up to 1 PFLOP FP4 / 1,000 TOPS sparse AI performance (manufacturer claim)

Model/workload scope: Vendor positions a single system for models up to 200B parameters and paired systems up to 405B

Best for: Local LLM inference · Fine-tuning · AI development · CUDA prototyping

128 GB is one CPU/GPU coherent memory pool, not separate 128 GB RAM plus 128 GB dedicated VRAM.

Official product page →

HP · Personal AI supercomputer

ZGX Nano G1n AI Station

Available
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x coherent unified system memory (273 GB/s)
AI engine
5th-generation Tensor Cores
Storage / OS
2 TB or 4 TB self-encrypted NVMe M.2 · NVIDIA DGX OS 7, Ubuntu 24.04

AI capability: Up to 1 PFLOP FP4 / 1,000 TOPS sparse AI performance (manufacturer claim)

Model/workload scope: Vendor positions a single system for models up to 200B parameters and paired systems up to 405B

Best for: Local LLM inference · Fine-tuning · AI development · CUDA prototyping

HP documents 128 GB coherent unified memory and no Windows support for this GB10 system.

Official product page →

Acer · Personal AI supercomputer

Veriton GN100 AI Mini Workstation

Regional availability
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x coherent unified system memory (273 GB/s)
AI engine
5th-generation Tensor Cores
Storage / OS
1 TB PCIe 4.0 or 4 TB NVMe M.2 configurations · NVIDIA DGX OS, Linux

AI capability: Up to 1 PFLOP FP4 / 1,000 TOPS sparse AI performance (manufacturer claim)

Model/workload scope: Designed for local prototyping, fine-tuning, testing and LLM deployment; paired systems are positioned for larger models

Best for: Local LLM inference · Fine-tuning · AI development · CUDA prototyping

128 GB is one CPU/GPU coherent memory pool, not separate 128 GB RAM plus 128 GB dedicated VRAM.

Official product page →

MSI · Personal AI supercomputer

EdgeXpert MS-C931

Available
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x coherent unified system memory (273 GB/s)
AI engine
5th-generation Tensor Cores
Storage / OS
1 TB or 4 TB self-encrypting NVMe M.2 · NVIDIA DGX OS, Linux

AI capability: Up to 1 PFLOP FP4 / 1,000 TOPS sparse AI performance (manufacturer claim)

Model/workload scope: Vendor positions a single system for models up to 200B parameters and paired systems up to 405B

Best for: Local LLM inference · Fine-tuning · AI development · CUDA prototyping

MSI explicitly labels the graphics memory as sharing the same 128 GB unified pool.

Official product page →

GIGABYTE · Personal AI supercomputer

AI TOP ATOM

Available
Compute platform
NVIDIA GB10 Grace Blackwell Superchip
Memory type
128 GB LPDDR5x coherent unified system memory (273 GB/s)
AI engine
5th-generation Tensor Cores
Storage / OS
1 TB PCIe 4.0 or 4 TB PCIe 4.0/5.0 configurations · NVIDIA DGX OS, Ubuntu Linux

AI capability: Up to 1 PFLOP FP4 / 1,000 TOPS sparse AI performance (manufacturer claim)

Model/workload scope: Vendor positions a single system for models up to 200B parameters and paired systems up to 405B

Best for: Local LLM inference · Fine-tuning · AI development · CUDA prototyping

128 GB is one CPU/GPU coherent memory pool, not separate 128 GB RAM plus 128 GB dedicated VRAM.

Official product page →

ASUS ROG · Portable AI 2-in-1

ROG Flow Z13 (2025) GZ302EA-XS99 (128 GB)

Regional availability
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X-8000 unified memory shared by CPU and GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
1 TB PCIe 4.0 NVMe SSD · Windows 11

AI capability: 50 NPU TOPS; GPU compute uses Radeon 8060S unified memory

Model/workload scope: ASUS positions the 128 GB configuration for local models up to 70B parameters

Best for: Portable local LLM work · AI-assisted creation · Development · Mobile workstation use

128 GB is unified memory. ASUS lists the XS99 as a real 128 GB SKU; it is not fixed 128 GB dedicated VRAM.

Official product page →

HP · Mobile AI workstation

ZBook Ultra G1a (128 GB)

Available
Compute platform
AMD Ryzen AI Max+ PRO 395
Memory type
128 GB LPDDR5X-8533 onboard unified memory; up to 96 GB GPU-addressable
AI engine
AMD Ryzen AI NPU, 50 TOPS
Storage / OS
Up to 4 TB PCIe NVMe SSD · Windows 11 Pro, Ubuntu Linux 24.04 (select configurations)

AI capability: 50 NPU TOPS; Radeon 8060S integrated GPU for accelerated local workloads

Model/workload scope: Designed for complex local AI and data-science workflows within unified-memory limits

Best for: Mobile AI workstation · Data science · Local inference · Professional creation

Memory is shared system memory on the Ryzen AI Max platform, not a discrete 128 GB VRAM pool.

Official product page →

Framework · Compact AI desktop

Framework Desktop Ryzen AI Max+ 395 (128 GB)

Available
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X-8000 unified memory; up to 96 GB graphics-addressable on Windows
AI engine
AMD XDNA NPU, up to 50 TOPS
Storage / OS
User-configurable NVMe storage · Windows 11, Linux

AI capability: Radeon 8060S GPU plus 50 TOPS-class NPU

Model/workload scope: Framework demonstrates local Llama 70B-class workloads on the 128 GB configuration

Best for: Local LLM inference · Linux AI development · Compact workstation · Upgradeable desktop

Framework explicitly distinguishes 128 GB total unified memory from up to 96 GB graphics-addressable memory.

Official product page →

GMKtec · Compact AI desktop

EVO-X2 Ryzen AI Max+ 395 (128 GB)

Available
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
128 GB configuration with dual M.2 slots; manufacturer states up to 16 TB total SSD capacity · Windows 11, Linux

AI capability: Ryzen AI Max platform with Radeon 8060S GPU and 50 TOPS NPU

Model/workload scope: Large local LLM inference and AI creation within 128 GB unified-memory and software-support limits

Best for: Local LLM inference · AI creation · Compact development workstation

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

MINISFORUM · Desktop AI workstation

MS-S1 MAX (128 GB)

Available
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
Dual M.2 slots with up to 16 TB total and RAID 0/1 support · Windows 11, Linux

AI capability: Up to 126 total platform TOPS (manufacturer claim); 50 TOPS NPU

Model/workload scope: MINISFORUM positions the 128 GB model for local 128B+ LLM workloads and multi-unit clustering

Best for: Local large-model inference · Rackable AI workstation · Multi-unit AI cluster

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

HP · Desktop AI workstation

Z2 Mini G1a Workstation (128 GB)

Available
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
1 TB or 2 TB NVMe SSD configurations · Windows 11 Pro

AI capability: 50 TOPS NPU plus Radeon 8060S integrated GPU acceleration

Model/workload scope: HP positions the system for local LLMs, AI-assisted professional work, CAD and content workflows

Best for: Professional local AI · CAD and visualization · Data analysis · Compact workstation fleets

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

CORSAIR · Desktop AI workstation

AI Workstation 300 (128 GB)

Regional availability
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
1 TB or 4 TB NVMe configurations · Windows 11 Home

AI capability: Up to 126 total platform TOPS; vendor software bundle for local AI workflows

Model/workload scope: CORSAIR lists GPT-OSS 120B, Qwen-class models and private local agents as target workloads

Best for: Preconfigured local AI · Private agents · Creative AI workflows

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

GEEKOM · Compact AI desktop

A9 Mega (128 GB)

Regional availability
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
2 TB included; dual M.2 expansion up to 8 TB · Windows 11 Pro, Linux-ready

AI capability: Up to 126 total platform TOPS (manufacturer claim); 50+ TOPS NPU

Model/workload scope: GEEKOM positions the system for Qwen, Llama, Stable Diffusion and DeepSeek-R1 70B local workflows

Best for: Local LLM inference · Stable Diffusion · Compact AI development · Creative workloads

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

Apple · Desktop AI workstation

Mac Studio with M3 Ultra (512 GB)

Previous generationSuperseded
Compute platform
Apple M3 Ultra
Memory type
Up to 512 GB unified memory shared across CPU, GPU and Neural Engine
AI engine
32-core Neural Engine
Storage / OS
Up to 16 TB SSD · macOS

AI capability: Apple GPU, Neural Engine and unified memory; no CUDA support

Model/workload scope: Very large local model inference where macOS/Metal software support is available

Best for: Large local inference · MLX workflows · Creative AI · macOS development

512 GB is unified memory, not dedicated VRAM; CUDA-only software is not compatible.

AI model compatibility report

Planning guidance only: memory fit and general framework support are separate. The exact model artifact, runtime version, context and workload must be tested before claiming it runs; no speed benchmark is implied.

Conservative usable memory: 409 GB

7–8B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–8B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Context length and KV cache consume memory beyond model weights.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

14B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 14B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Tokenizer, context and quantization format change actual memory use.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

30–35B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 30–35B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Long context can move a memory-fit model into swap or failure.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

70B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 70B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

This is a memory-fit estimate, not a speed claim. Context and runtime overhead matter.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

100B+ dense or MoE model, 4-bit class

Comfortable memory fitExperimental support

Comfortable memory headroom for 100B+ dense or MoE model, 4-bit class; speed has not been measured. Software support is experimental and depends on the model and framework version.

MoE architectures vary widely; verify the exact artifact and active-expert implementation.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · class-v1

SDXL-class image generation

Comfortable memory fitFramework supported

Comfortable memory headroom for SDXL-class image generation; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Resolution, ControlNet, batch size and app implementation change peak memory.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · General framework documentation · sdxl-class-v1

FLUX-class quantized image generation

Comfortable memory fitExperimental support

Comfortable memory headroom for FLUX-class quantized image generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Compatibility depends on the exact model build and framework; CUDA-only nodes remain unavailable.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · flux-class-v1

Whisper large-class transcription

Comfortable memory fitFramework supported

Comfortable memory headroom for Whisper large-class transcription; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Language, batch size and implementation affect throughput.

Frameworks: MLX Whisper · Core ML builds · PyTorch MPS

Evidence: Memory estimate · General framework documentation · whisper-large-class-v1

7–14B LoRA / QLoRA fine-tuning

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–14B LoRA / QLoRA fine-tuning; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Dataset, optimizer states, sequence length and rank can raise memory use substantially.

Frameworks: MLX · PyTorch MPS when operations are supported

Evidence: Memory estimate · General framework documentation · lora-class-v1

Local diffusion video generation

Comfortable memory fitExperimental support

Comfortable memory headroom for Local diffusion video generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Many video pipelines and custom nodes require CUDA; memory fit alone does not prove compatibility.

Frameworks: Model-specific Metal/Core ML/MLX ports

Evidence: Memory estimate · General framework documentation · video-class-v1

CUDA-only models and custom nodes

Comfortable memory fitUnsupported

Comfortable memory headroom for CUDA-only models and custom nodes; speed has not been measured. The required software ecosystem is unsupported on this platform.

Apple Silicon does not support CUDA. Use a Metal/MLX/Core ML alternative or NVIDIA hardware.

Frameworks: NVIDIA CUDA

Evidence: General framework documentation · cuda-ecosystem-v1

Official product page →

NVIDIA · Desktop AI workstation

DGX Station with GB300

Announced
Compute platform
NVIDIA GB300 Grace Blackwell Desktop Superchip
Memory type
Up to 748 GB coherent memory: up to 252 GB GPU-local HBM3e + 496 GB CPU LPDDR5X
AI engine
NVIDIA Tensor Cores and NVLink-C2C coherent interconnect
Storage / OS
NVMe storage configuration varies by DGX Station offering · NVIDIA DGX OS 7 (Ubuntu-based)

AI capability: GB300 Grace Blackwell Desktop Superchip; no normalized consumer TOPS comparison is claimed

Model/workload scope: NVIDIA positions the system for local models up to 1T parameters and two-system scaling for larger workloads

Best for: Trillion-parameter-class local experimentation · Enterprise AI development · Fine-tuning · Large inference

The 252 GB HBM3e and 496 GB LPDDR5X form one coherent address space; they must not be represented as 748 GB dedicated VRAM.

Official product page →

Dell · Desktop AI workstation

Pro Max Tower T2 with RTX PRO 6000 Blackwell

Available
Compute platform
Up to Intel Core Ultra 9 285K + 1× NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Memory type
128 GB system memory + 1×96 GB dedicated VRAM (96 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Configurable NVMe and SATA storage · Windows 11 Pro, Windows 11 Pro for Workstations, Ubuntu Linux 24.04 LTS, Red Hat Enterprise Linux 9.6

AI capability: Up to 1 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: Local 70B-class inference · AI development · CAD · Professional visualization

96 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 96 GB allocation.

Official product page →

HP · Desktop AI workstation

Z8 Fury G6i with 4× RTX PRO 6000 Blackwell Max-Q

Available
Compute platform
Up to Intel Xeon 600 series + 4× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Memory type
2048 GB system memory + 4×96 GB dedicated VRAM (384 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Large configurable professional NVMe/SATA storage · Windows 11 Pro for Workstations, Linux options vary by configuration

AI capability: Up to 4 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: Multi-GPU AI development · Simulation · VFX · Shared workstation compute

384 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 384 GB allocation.

Official product page →

Lenovo · Desktop AI workstation

ThinkStation PX with 4× RTX PRO 6000 Blackwell Max-Q

Regional availability
Compute platform
Up to dual Intel Xeon Scalable processors + 4× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Memory type
2048 GB system memory + 4×96 GB dedicated VRAM (384 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Up to seven or more drives depending on bay configuration · Windows 11 Pro for Workstations, Red Hat Enterprise Linux certification

AI capability: Up to 4 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: Multi-GPU AI · Rendering · Simulation · Enterprise workstation workloads

384 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 384 GB allocation.

Official product page →

BOXX · Desktop AI workstation

APEXX T4 PRO-X with 4× RTX PRO 6000 Blackwell Max-Q

Available
Compute platform
AMD Ryzen Threadripper PRO 9000 WX-series (up to 96 cores) + 4× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Memory type
2048 GB system memory + 4×96 GB dedicated VRAM (384 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Multiple NVMe, 2.5-inch and 3.5-inch drive options · Windows 11 Pro for Workstations, Linux by configuration

AI capability: Up to 4 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: AI model training · Machine learning development · Rendering · Complex simulation

384 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 384 GB allocation.

Official product page →

AMD · Compact AI desktop

Ryzen AI Halo Developer Platform (128 GB)

Regional availability
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
2 TB M.2 self-encrypting SSD · Linux, Windows 11

AI capability: Radeon 8060S plus 50 TOPS XDNA 2 NPU; 256 GB/s memory bandwidth

Model/workload scope: AMD positions the system for local models up to 200B parameters

Best for: ROCm development · Local LLM inference · AI agents · Cross-platform AI development

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

BOSGAME · Compact AI desktop

M5 AI Mini Desktop (128 GB)

Available
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
2 TB PCIe 4.0 NVMe SSD with dual M.2 slots · Windows

AI capability: Up to 126 total platform TOPS and 50 TOPS NPU (manufacturer claim)

Model/workload scope: Vendor positions the 128 GB SKU for private local LLM workflows

Best for: Local LLM inference · Creative AI · Compact AI workstation

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

ACEMAGIC · Compact AI desktop

F9A Ryzen AI Max+ 395 Mini PC (128 GB)

Regional availability
Compute platform
AMD Ryzen AI Max+ 395
Memory type
128 GB LPDDR5X unified memory shared by CPU and Radeon 8060S GPU
AI engine
AMD XDNA NPU, 50 TOPS
Storage / OS
Dual M.2 PCIe 4.0 x4 NVMe slots · Windows

AI capability: Up to 126 total platform TOPS and 50 TOPS NPU (manufacturer claim)

Model/workload scope: Vendor positions the system for local inference of models up to 120B parameters

Best for: Local LLM inference · AI development · Content creation

128 GB is unified system/GPU memory. The Radeon 8060S has no separate 128 GB dedicated VRAM pool.

Official product page →

Dell · Desktop AI workstation

Pro Max with GB300 (FCT6263)

Available
Compute platform
NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip
Memory type
496 GB CPU LPDDR5X + 252 GB GPU-local HBM3e (748 GB coherent address space)
AI engine
NVIDIA Tensor Cores and NVLink-C2C coherent interconnect
Storage / OS
Up to four 4 TB PCIe Gen4 self-encrypting NVMe SSDs · Ubuntu 24.04 LTS with NVIDIA AI Developer Tools

AI capability: Up to 20 PFLOPS FP4 AI performance (manufacturer claim)

Model/workload scope: Vendor positions the platform for local models up to 1T parameters

Best for: Autonomous AI agents · Large local inference · Fine-tuning · Enterprise AI development

748 GB is a coherent address space made of 496 GB LPDDR5X and 252 GB GPU-local HBM3e; it is not 748 GB dedicated VRAM.

Official product page →

ASUS · Desktop AI workstation

ExpertCenter Pro ET900N G3

Available
Compute platform
NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip
Memory type
496 GB CPU LPDDR5X + 252 GB GPU-local HBM3e (748 GB coherent address space)
AI engine
NVIDIA Tensor Cores and NVLink-C2C coherent interconnect
Storage / OS
Two preinstalled 2 TB NVMe SSDs in RAID 1 plus two additional M.2 slots · NVIDIA DGX Station software architecture, Windows support planned

AI capability: Up to 20 PFLOPS FP4 AI performance (manufacturer claim)

Model/workload scope: Vendor positions the platform for local models up to 1T parameters

Best for: Trillion-parameter local inference · AI agents · Enterprise AI development

748 GB is a coherent address space made of 496 GB LPDDR5X and 252 GB GPU-local HBM3e; it is not 748 GB dedicated VRAM.

Official product page →

HP · Desktop AI workstation

ZGX Fury AI Station

Announced
Compute platform
NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip
Memory type
496 GB CPU LPDDR5X + 252 GB GPU-local HBM3e (748 GB coherent address space)
AI engine
NVIDIA Tensor Cores and NVLink-C2C coherent interconnect
Storage / OS
Storage configuration not yet published by HP · Ubuntu with NVIDIA AI Developer Tools, Windows support planned

AI capability: Up to 20 PFLOPS FP4 AI performance (manufacturer claim)

Model/workload scope: Vendor positions the platform for local models up to 1T parameters

Best for: Department-scale AI inference · Concurrent local AI users · Long-running agents

748 GB is a coherent address space made of 496 GB LPDDR5X and 252 GB GPU-local HBM3e; it is not 748 GB dedicated VRAM.

Official product page →

Apple · Desktop AI workstation

Mac Studio with M4 Max (128 GB)

Previous generationSuperseded
Compute platform
Apple M4 Max
Memory type
128 GB unified memory shared across CPU, GPU and Neural Engine (up to 546 GB/s)
AI engine
16-core Neural Engine
Storage / OS
Up to 8 TB SSD · macOS

AI capability: Apple GPU and Neural Engine acceleration; no CUDA support

Model/workload scope: Local inference where macOS Metal or MLX software support is available

Best for: MLX workflows · Creative AI · macOS development · Local inference

128 GB is unified system memory, not dedicated VRAM; CUDA-only software is not compatible.

AI model compatibility report

Planning guidance only: memory fit and general framework support are separate. The exact model artifact, runtime version, context and workload must be tested before claiming it runs; no speed benchmark is implied.

Conservative usable memory: 102 GB

7–8B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–8B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Context length and KV cache consume memory beyond model weights.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

14B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 14B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Tokenizer, context and quantization format change actual memory use.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

30–35B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 30–35B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Long context can move a memory-fit model into swap or failure.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

70B LLM, 4-bit

Comfortable memory fitFramework supported

Comfortable memory headroom for 70B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

This is a memory-fit estimate, not a speed claim. Context and runtime overhead matter.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · General framework documentation · class-v1

100B+ dense or MoE model, 4-bit class

Expected to fitExperimental support

100B+ dense or MoE model, 4-bit class is expected to fit; speed has not been measured. Software support is experimental and depends on the model and framework version.

MoE architectures vary widely; verify the exact artifact and active-expert implementation.

Frameworks: MLX · llama.cpp / LM Studio

Evidence: Memory estimate · class-v1

SDXL-class image generation

Comfortable memory fitFramework supported

Comfortable memory headroom for SDXL-class image generation; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Resolution, ControlNet, batch size and app implementation change peak memory.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · General framework documentation · sdxl-class-v1

FLUX-class quantized image generation

Comfortable memory fitExperimental support

Comfortable memory headroom for FLUX-class quantized image generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Compatibility depends on the exact model build and framework; CUDA-only nodes remain unavailable.

Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported

Evidence: Memory estimate · flux-class-v1

Whisper large-class transcription

Comfortable memory fitFramework supported

Comfortable memory headroom for Whisper large-class transcription; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Language, batch size and implementation affect throughput.

Frameworks: MLX Whisper · Core ML builds · PyTorch MPS

Evidence: Memory estimate · General framework documentation · whisper-large-class-v1

7–14B LoRA / QLoRA fine-tuning

Comfortable memory fitFramework supported

Comfortable memory headroom for 7–14B LoRA / QLoRA fine-tuning; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.

Dataset, optimizer states, sequence length and rank can raise memory use substantially.

Frameworks: MLX · PyTorch MPS when operations are supported

Evidence: Memory estimate · General framework documentation · lora-class-v1

Local diffusion video generation

Comfortable memory fitExperimental support

Comfortable memory headroom for Local diffusion video generation; speed has not been measured. Software support is experimental and depends on the model and framework version.

Many video pipelines and custom nodes require CUDA; memory fit alone does not prove compatibility.

Frameworks: Model-specific Metal/Core ML/MLX ports

Evidence: Memory estimate · General framework documentation · video-class-v1

CUDA-only models and custom nodes

Comfortable memory fitUnsupported

Comfortable memory headroom for CUDA-only models and custom nodes; speed has not been measured. The required software ecosystem is unsupported on this platform.

Apple Silicon does not support CUDA. Use a Metal/MLX/Core ML alternative or NVIDIA hardware.

Frameworks: NVIDIA CUDA

Evidence: General framework documentation · cuda-ecosystem-v1

Official product page →

HP · Desktop AI workstation

Z2 Tower G1i with RTX PRO 6000 Blackwell

Available
Compute platform
Up to Intel Core Ultra 9 285K + 1× NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Memory type
256 GB system memory + 1×96 GB dedicated VRAM (96 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Up to 36 TB total storage · Windows 11 Pro, Linux options

AI capability: Up to 1 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: Local AI inference · AI development · CAD · Professional visualization

96 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 96 GB allocation.

Official product page →

HP · Desktop AI workstation

Z8 G5 with 2× RTX PRO 6000 Blackwell Max-Q

Available
Compute platform
Up to dual Intel Xeon processors + 2× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Memory type
1024 GB system memory + 2×96 GB dedicated VRAM (192 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Up to 136 TB total storage · Windows 11 Pro for Workstations, Linux options

AI capability: Up to 2 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: Multi-GPU AI · Model training · Rendering · Simulation

192 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 192 GB allocation.

Official product page →

HP · Desktop AI workstation

Z6 G5 A with 3× RTX PRO 6000 Blackwell Max-Q

Available
Compute platform
Up to AMD Ryzen Threadripper PRO 9000 WX-series + 3× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Memory type
1024 GB system memory + 3×96 GB dedicated VRAM (288 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Up to 88 TB total storage · Windows 11 Pro for Workstations, Linux options

AI capability: Up to 3 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: Multi-GPU AI · Model training · VFX · Virtual production

288 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 288 GB allocation.

Official product page →

Lenovo · Desktop AI workstation

ThinkStation P5 Gen 2 with 2× RTX PRO 6000 Blackwell Max-Q

Regional availability
Compute platform
Up to Intel Xeon 600 series + 2× NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Memory type
1024 GB system memory + 2×96 GB dedicated VRAM (192 GB aggregate across GPUs)
AI engine
NVIDIA Blackwell 5th-generation Tensor Cores
Storage / OS
Multiple M.2 SSD and 3.5-inch drive options · Windows 11 Pro for Workstations, Linux options

AI capability: Up to 2 RTX PRO 6000 Blackwell GPU configuration; no normalized whole-system benchmark is claimed

Model/workload scope: Model capacity depends on precision, framework and whether the workload supports tensor/model parallelism across separate GPUs

Best for: AI development · Rendering · Simulation · Professional visualization

192 GB is the sum of separate 96 GB GPU memories. It is not automatically one unified 192 GB allocation.

Official product page →

Need CUDA GPUs or a custom multi-GPU build?

Ready unified-memory systems are not always the right answer. CUDA-only software, multi-GPU training, replaceable components or verified local retailer pricing may make a custom workstation the better choice.

Build a custom AI workstation