Running a 7B-parameter language model at 70 tokens per second on a single-board computer sounds aspirational, but Forlinx Embedded is now listing an M.2 2280 AI accelerator card built on Rockchip's RK1820 and RK1828 coprocessors that hits exactly those numbers. The card slots into a standard M-Key M.2 socket and offloads AI inference entirely from the host processor, letting boards based on the RK3568, RK3572, RK3576, and RK3588 run large language models, vision-language models, and conventional neural networks without cloud connectivity. The RK1828 variant packs 5GB of 3D-stacked in-package DRAM alongside a 20 TOPS (INT8) NPU, while the RK1820 ships with 2.5GB. That stacked memory architecture delivers high internal bandwidth for transformer workloads and avoids competing with the host board's system RAM.

The software side leans on Rockchip's RKNN3 SDK, which handles model conversion, quantization, and deployment for CNN, LLM, and VLM workloads. It accepts models from TensorFlow, PyTorch, and ONNX, and runs on Linux and Android hosts. The NPU supports INT4, INT8, INT16, FP8, FP16, and BF16 precision formats, giving developers flexibility to trade accuracy for throughput depending on the workload. Forlinx's own benchmarks on an OK3588-C development board show roughly 102 tokens per second for Qwen2.5-3B and 70 tokens per second for Qwen2.5-7B with a single RK1828 card, measured at 128 input and 128 output tokens in performance mode.

Forlinx states it has completed in-depth driver debugging and full operator-implementation verification for the RK182X series on both Linux and Android hosts. Community interest in the underlying RK1820/RK1828 platform is already visible outside Forlinx's own documentation: an independently maintained setup guide on GitHub walks through converting a DeepSeek-R1-Distill-Qwen-1.5B model to Rockchip's .rkllm format and running it on the NPU of a Firefly RK182X development kit, evidence that hobbyists are already picking up the RKNN3 workflow on hardware built around the same coprocessors as Forlinx's card.

What sets this apart from other M.2 AI accelerators like the Hailo-8 (26 TOPS, but limited to vision workloads) or Google's largely end-of-life Coral TPU (4 TOPS) is the integrated high-bandwidth memory that makes on-device LLM inference practical in this form factor. Forlinx has also demonstrated a four-card PCIe cascade configuration on the OK3588-C, using pipeline parallelism to distribute transformer layers across the accelerators. That setup handles 27B and 31B parameter models at roughly 13 tokens per second with a combined accelerator power draw of about 40W, and Rockchip's next-generation RK1860 coprocessor is expected later in 2026 with further improvements to that ceiling.

Forlinx has not published official retail pricing. The card communicates over a single PCIe 2.1 lane.