Running a 7B-parameter language model at 70 tokens per second from a standard M.2 slot is now within reach for Rockchip-based embedded Linux boards. Forlinx Embedded has released an M.2 2280 AI accelerator built around Rockchip's RK1820 and RK1828 coprocessors, each delivering 20 TOPS of INT8 NPU performance with 3D-stacked DRAM integrated directly into the package. The RK1820 packs 2.5GB of on-package memory for models up to 3B parameters, while the RK1828 doubles that to 5GB for 7B-class workloads. Both cards connect to the host over a single PCIe 2.1 lane and support INT4, INT8, INT16, FP8, FP16, and BF16 precision formats.

The on-package memory architecture is what makes these cards particularly interesting for LLM inference, where bandwidth between processor and memory typically limits token throughput. In Forlinx's benchmarks using an RK1828 on an OK3588-C development board, the card achieved approximately 102 tokens per second on Qwen2.5-3B and 70 tokens per second on Qwen2.5-7B, measured with 128-token input and output sequences. Compatible host processors include the RK3568, RK3572, RK3576, and RK3588, covering a broad swath of existing embedded boards from manufacturers like Firefly, Radxa, and Pine64.

Forlinx has also demonstrated a four-card PCIe cascade that distributes Transformer layers across accelerators using pipeline parallelism. With four RK1828 modules and their combined 20GB of on-package memory, the setup handles 27B to 31B parameter models at roughly 13 tokens per second, drawing approximately 40W across all four cards. That makes fully offline inference with models like Qwen2.5-31B feasible on an embedded platform, no cloud connectivity required.

On the software side, the cards use Rockchip's RKNN3 toolkit for model conversion, quantization, and deployment. RKNN3 is a distinct SDK generation from Rockchip's earlier RKNN Toolkit 2 line, built specifically around the RK1820, RK1828, and RK3572, and Rockchip maintains it as open repositories on GitHub: rknn3-toolkit for PC-side model conversion and inference, and rknn3-model-zoo for reference CNN, LLM, and VLM deployment examples, alongside prebuilt firmware for RK1820/RK1828 development boards. Rockchip formally released RKNN3 SDK V1.0.0 in June 2026. Separately, Rockchip also publishes the RKNN Toolkit 2 and RKNN-LLM repositories on GitHub for its other SoC lines, and the NPU kernel driver source is available in the Rockchip kernel tree. The toolchain accepts models from TensorFlow, PyTorch, and ONNX, with Linux and Android as supported host operating systems.

Standalone M.2 module pricing has not been publicly listed by Forlinx, though development kits bundling an RK1828 with an RK3588 carrier board have been available starting at roughly $1,030 (€950). DFRobot also offers the RK1828 in M.2 form with setup guides for RK3576-based hosts. A community-written guide on GitHub, rk182x-evk-setup-guide, walks through configuring a Firefly development kit built around the same RK1820/RK1828 coprocessors, from unboxing through running a local LLM on the NPU. It targets the SO-DIMM devkit form factor rather than Forlinx's M.2 card specifically, but it exercises the same RKNN3 software stack. Rockchip's next-generation AI coprocessor, the RK1860, sits on the company's public roadmap as the successor with higher performance targets.