Running 284-billion-parameter language models locally just became a practical proposition on the desktop. Minisforum has started selling the MS-S1 Max-P495, a compact workstation built around AMD's flagship Ryzen AI Max+ PRO 495 "Gorgon Halo" APU. The system pairs 192GB of LPDDR5x-8533 memory with the ability to allocate up to 160GB as VRAM, a unified memory pool large enough to run massive models like DeepSeek V4 Flash entirely on-device without a discrete GPU.
The Gorgon Halo APU steps up from the Strix Halo silicon found in the original MS-S1 Max, bringing 16 Zen 5 cores and 32 threads with boost clocks up to 5.2 GHz, a 40-CU Radeon 8065S iGPU on RDNA 3.5 architecture, and a 55 TOPS XDNA 2 NPU. Memory bandwidth improves as well, with LPDDR5x support climbing to 8,533 MT/s from the previous generation's ceiling. For storage and expansion, the system includes two M.2 slots supporting up to 8TB SSDs each, a PCIe x16 slot wired for PCIe 4.0 x4 bandwidth, two USB4 v2 ports, and two standard USB4 ports.
Linux users eyeing the MS-S1 Max-P495 for local inference work should be aware that unlocking the full VRAM allocation on AMD APUs under Linux requires tuning kernel parameters rather than relying on BIOS settings alone. The default dedicated VRAM allocation is typically small, and the system relies on the kernel's Graphics Translation Table to map system memory for GPU use. Adjusting amdgpu.gttsize and TTM page limits through GRUB is the standard approach to making the full 160GB pool available for inference workloads. AMD's ROCm documentation covers optimization for these APUs, and the Gorgon Halo platform follows the same patterns established by Strix Halo.
On the software side, AMD's ROCm 7.14 added Gorgon Halo support in July 2026, and ROCm identifies these integrated GPUs as gfx1151 devices, the same target used by Strix Halo. Community projects for local inference on that target include kyuz0's amd-strix-halo-toolboxes, a set of prebuilt llama.cpp containers with Vulkan and ROCm backends for Strix Halo systems. Benchmarks on Strix Halo hardware show neither backend is fastest across the board, with ROCm ahead on prompt processing and Vulkan ahead on token generation. Those results come from Strix Halo machines rather than the P495. The original MS-S1 Max also has a community-maintained guide and script for updating its BIOS from Linux, since Minisforum ships its update tooling for Windows.
The MS-S1 Max-P495 ships in a single configuration with 192GB of RAM and a 2TB SSD for $7,399 (€6,800). The non-P495 MS-S1 Max, which tops out at 128GB of RAM and 96GB of allocatable VRAM on the Strix Halo APU, starts at $3,799 (€3,500) for those who don't need the extra memory headroom.



