Planning guidance only: memory fit and general framework support are separate. The exact model artifact, runtime version, context and workload must be tested before claiming it runs; no speed benchmark is implied.
7–8B LLM, 4-bit
Comfortable memory fitFramework supported
Comfortable memory headroom for 7–8B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.
Context length and KV cache consume memory beyond model weights.
Frameworks: MLX · llama.cpp / LM Studio
Evidence: Memory estimate · General framework documentation · class-v1
14B LLM, 4-bit
Comfortable memory fitFramework supported
Comfortable memory headroom for 14B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.
Tokenizer, context and quantization format change actual memory use.
Frameworks: MLX · llama.cpp / LM Studio
Evidence: Memory estimate · General framework documentation · class-v1
30–35B LLM, 4-bit
Comfortable memory fitFramework supported
Comfortable memory headroom for 30–35B LLM, 4-bit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.
Long context can move a memory-fit model into swap or failure.
Frameworks: MLX · llama.cpp / LM Studio
Evidence: Memory estimate · General framework documentation · class-v1
70B LLM, 4-bit
Tight memory fitFramework supported
70B LLM, 4-bit is a tight memory fit; keep context, batch size and related settings conservative. General framework support exists; the exact model artifact has not been verified on this device.
This is a memory-fit estimate, not a speed claim. Context and runtime overhead matter.
Frameworks: MLX · llama.cpp / LM Studio
Evidence: Memory estimate · General framework documentation · class-v1
100B+ dense or MoE model, 4-bit class
Insufficient safe memoryExperimental support
There is no safe memory headroom for 100B+ dense or MoE model, 4-bit class. Software support is experimental and depends on the model and framework version.
MoE architectures vary widely; verify the exact artifact and active-expert implementation.
Frameworks: MLX · llama.cpp / LM Studio
Evidence: Memory estimate · class-v1
SDXL-class image generation
Comfortable memory fitFramework supported
Comfortable memory headroom for SDXL-class image generation; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.
Resolution, ControlNet, batch size and app implementation change peak memory.
Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported
Evidence: Memory estimate · General framework documentation · sdxl-class-v1
FLUX-class quantized image generation
Comfortable memory fitExperimental support
Comfortable memory headroom for FLUX-class quantized image generation; speed has not been measured. Software support is experimental and depends on the model and framework version.
Compatibility depends on the exact model build and framework; CUDA-only nodes remain unavailable.
Frameworks: Draw Things · MLX/Core ML builds · PyTorch MPS when supported
Evidence: Memory estimate · flux-class-v1
Whisper large-class transcription
Comfortable memory fitFramework supported
Comfortable memory headroom for Whisper large-class transcription; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.
Language, batch size and implementation affect throughput.
Frameworks: MLX Whisper · Core ML builds · PyTorch MPS
Evidence: Memory estimate · General framework documentation · whisper-large-class-v1
7–14B LoRA / QLoRA fine-tuning
Expected to fitFramework supported
7–14B LoRA / QLoRA fine-tuning is expected to fit; speed has not been measured. General framework support exists; the exact model artifact has not been verified on this device.
Dataset, optimizer states, sequence length and rank can raise memory use substantially.
Frameworks: MLX · PyTorch MPS when operations are supported
Evidence: Memory estimate · General framework documentation · lora-class-v1
Local diffusion video generation
Tight memory fitExperimental support
Local diffusion video generation is a tight memory fit; keep context, batch size and related settings conservative. Software support is experimental and depends on the model and framework version.
Many video pipelines and custom nodes require CUDA; memory fit alone does not prove compatibility.
Frameworks: Model-specific Metal/Core ML/MLX ports
Evidence: Memory estimate · General framework documentation · video-class-v1
CUDA-only models and custom nodes
Comfortable memory fitUnsupported
Comfortable memory headroom for CUDA-only models and custom nodes; speed has not been measured. The required software ecosystem is unsupported on this platform.
Apple Silicon does not support CUDA. Use a Metal/MLX/Core ML alternative or NVIDIA hardware.
Frameworks: NVIDIA CUDA
Evidence: General framework documentation · cuda-ecosystem-v1