Forlinx Launched M.2 AI Accelerator Card
The new hardware integrates Rockchip processors to support offline inference for small language models.
Updated on Sept. 29, 2026 in Semiconductors

Live Poll
Would you prefer to run AI models on your own hardware instead of using cloud services?
Forlinx Embedded has introduced a new M.2 AI accelerator card powered by Rockchip RK1820 and RK1828 processors. The device is designed to handle offline inference for language models ranging from 3B to 7B parameters.
Why it matters
The card features a 3D stacked architecture that integrates DRAM directly with the processor to overcome memory bandwidth constraints. This design reduces reliance on host system memory, enabling localized processing of language models.
The accelerator supports INT4, INT8, INT16, FP8, FP16, and BF16 formats and delivers 20 TOPS of computing power. It uses an M.2 2280 M-Key form factor and interfaces via a single PCIe 2.1 lane.
The players
Forlinx Embedded
This company specializes in the development and manufacturing of embedded computing hardware and modules.
Rockchip
This semiconductor company develops system-on-chip solutions for mobile and embedded internet of things devices.
The details
The system utilizes pipeline parallelism to distribute Transformer layers across multiple cascaded cards, allowing for the execution of models as large as 27B parameters. Compatible with Linux and Android, the hardware supports TensorFlow, PyTorch, and ONNX frameworks using the Rockchip RKNN3 SDK.
Timeline
September 29, 2026: The product release was finalized.
The Tech Race
This release reflects the broader industry move to integrate AI inference at the edge using standardized form factors like the M.2 2280. By placing memory in the same package as the processor, Forlinx challenges the dominance of host-heavy memory architectures.
Developers and systems integrators can use this card to add offline AI capabilities to standard devices without requiring massive system memory upgrades. The hardware allows for specific throughputs like 102 tokens per second for Qwen2.5-3B models.
The takeaway
This development highlights the trend toward localized AI processing in compact, power-efficient form factors. Integrating DRAM directly into the processor package represents a practical solution for overcoming bandwidth limitations in edge hardware.
Further reading
For more on how new architectures are changing performance, see our Semiconductors section.
Source note: This article includes information reported by LinuxGizmos.
Live Poll
Would you prefer to run AI models on your own hardware instead of using cloud services?







