Forlinx Launched M.2 AI Accelerator Card

The new hardware integrates Rockchip processors to support offline inference for small language models.

Updated on Sept. 29, 2026 in Semiconductors

Isometric editorial illustration of a generic M.2 hardware accelerator card with stacked silicon components and gold interface pins.
Forlinx Embedded has introduced a new M.2 AI accelerator card utilizing Rockchip processors to facilitate offline inference for language models. AI Illustration. Upload story photo >

Live Poll

Would you prefer to run AI models on your own hardware instead of using cloud services?

Forlinx Embedded has introduced a new M.2 AI accelerator card powered by Rockchip RK1820 and RK1828 processors. The device is designed to handle offline inference for language models ranging from 3B to 7B parameters.

Why it matters

The card features a 3D stacked architecture that integrates DRAM directly with the processor to overcome memory bandwidth constraints. This design reduces reliance on host system memory, enabling localized processing of language models.

The accelerator supports INT4, INT8, INT16, FP8, FP16, and BF16 formats and delivers 20 TOPS of computing power. It uses an M.2 2280 M-Key form factor and interfaces via a single PCIe 2.1 lane.

The players

Forlinx Embedded

This company specializes in the development and manufacturing of embedded computing hardware and modules.

Rockchip

This semiconductor company develops system-on-chip solutions for mobile and embedded internet of things devices.

The details

The system utilizes pipeline parallelism to distribute Transformer layers across multiple cascaded cards, allowing for the execution of models as large as 27B parameters. Compatible with Linux and Android, the hardware supports TensorFlow, PyTorch, and ONNX frameworks using the Rockchip RKNN3 SDK.

Timeline

  1. September 29, 2026: The product release was finalized.

The Tech Race

This release reflects the broader industry move to integrate AI inference at the edge using standardized form factors like the M.2 2280. By placing memory in the same package as the processor, Forlinx challenges the dominance of host-heavy memory architectures.

Developers and systems integrators can use this card to add offline AI capabilities to standard devices without requiring massive system memory upgrades. The hardware allows for specific throughputs like 102 tokens per second for Qwen2.5-3B models.

The takeaway

This development highlights the trend toward localized AI processing in compact, power-efficient form factors. Integrating DRAM directly into the processor package represents a practical solution for overcoming bandwidth limitations in edge hardware.

Further reading

For more on how new architectures are changing performance, see our Semiconductors section.

Source note: This article includes information reported by LinuxGizmos.

Live Poll

Would you prefer to run AI models on your own hardware instead of using cloud services?