We use some essential cookies to make our website work.

We use optional cookies, as detailed in our cookie policy, to remember your settings and understand how you use our website.

Seeing, understanding, and responding: low-power CNN, VLM and SLM workloads on Raspberry Pi 5

Most edge AI/ML workloads belong on the CPU, but accelerators are important for the small minority of workloads that the CPU can’t accommodate. Raspberry Pi users have a variety of choices for acceleration, with our own AI HAT+ and HAT+ 2 as well as offerings from the wider Raspberry Pi ecosystem. Here, our friends at Sixfab and DEEPX discuss running vision and compact language AI on Raspberry Pi 5 with the Sixfab AI HAT+, and what a three-watt NPU makes possible at the edge.

Raspberry Pi has made AI development accessible to millions of engineers, students and makers. The next step is moving from a model that runs once, in a demo, to an intelligent system that keeps watching, understanding and responding — without depending on a constant cloud connection.

The Sixfab AI HAT+ for Raspberry Pi 5 is a third-party AI accelerator board built around the DEEPX DX-M1M NPU. It adds 25 TOPS of dedicated AI acceleration at approximately 3 W of typical sustained NPU power, while leaving the Raspberry Pi’s CPU free for camera handling, application logic, connectivity and device control. In this article we’d like to show what that combination is genuinely good at, where it isn’t the right tool, and how to get from first boot to first inference in a few minutes.

Why three watts, not the TOPS number, is the headline

Peak TOPS figures make good marketing, but most edge products live and die by their power and thermal budget. The real system still needs headroom for the Raspberry Pi itself, a camera, storage, networking and whatever the device actually controls. An accelerator that stays around 3 W under sustained load means smaller enclosures, simpler (often passive) cooling, and AI that can run continuously rather than in short bursts.

That efficiency is what makes local AI practical in the places where it matters most: smart cameras, robots, industrial equipment and distributed sensors, where cloud latency, connectivity or data privacy would otherwise become the limiting constraint.

ENGINEERING NOTE:  The ~3 W figure refers to typical sustained NPU power for the 25 TOPS DX-M1M option, not complete system wall power. Always measure your final Raspberry Pi configuration with its actual camera, storage, workload and cooling.

One NPU for three classes of intelligence

Most edge products begin with perception: a CNN detects an object, segments a work area, estimates a pose or flags an abnormal event. The next generation of edge systems also needs to understand context and communicate what it sees. The DX-M1M is designed to support that whole journey on a single piece of silicon and one toolchain.

CNN: efficient visual perception

Object detection, classification, segmentation, pose estimation, depth and image enhancement turn camera pixels into structured events. The DEEPX ModelZoo ships pre-optimised examples — including YOLO-family detectors, segmentation and pose models — and DX-Stream helps you assemble them into GStreamer-based camera pipelines.

Compact VLM: connecting vision with meaning

A compact vision-language model adds contextual understanding on top of visual results: describing a scene, answering a focused question about an image, or interpreting an event in natural language — locally, with nothing leaving the device.

SLM: talking to the edge system

A small language model can interpret a user command, summarise local events or generate a concise explanation, creating a natural interface for cameras, robots and machines without routing every interaction through the cloud.

What about memory?

The first question any developer asks about language models at the edge is how much memory is available and how large a model will fit. The DX-M1M module carries 2 GB of dedicated on-module LPDDR4x, which comfortably supports VLM and SLM workloads alongside vision models.

Accuracy matters more than a peak benchmark

Edge AI performance is only useful if the model stays accurate on the images and conditions your application actually faces. DEEPX treats INT8 optimisation as an accuracy-aware engineering process, not a simple conversion step:

  1. Start with a known FP32 accuracy baseline and a clearly defined application metric.
  2. Calibrate with representative data from the real deployment environment — not only clean sample images.
  3. Compile and optimise the model with DX-COM into a portable .dxnn deployment artifact.
  4. Compare accuracy, latency, power and thermal behaviour on the final Raspberry Pi system.

This matters most in the difficult conditions edge cameras actually meet: low light, glare, motion blur, occlusion and unusual angles. A credible result shows both the efficiency gain and the retained application accuracy.

How fast is it, really?

Numbers beat adjectives. The figures below were measured on a Raspberry Pi 5 (8 GB) running the shipping Sixfab software release, comparing the DX-M1M and DX-M1 NPUs:

WorkloadDX-M1M (NPU)DX-M1 (NPU)
mobilenet_v2, 240×2402361 FPS3223 FPS
deeplabv3plus, 512×512155 FPS231 FPS
Qwen3-1.7B, 96 prefill tokensTTFT: 599.04 ms TPS: 4.64 tok/sTTFT: 544.58 ms TPS: 11.60 tok/s

* The DX-M1 uses LPDDR5, while the DX-M1M uses LPDDR4X, resulting in a performance difference due to their memory bandwidths.

From first boot to first inference

The fastest way to understand the platform is to verify the hardware, run a packaged model and watch the NPU work. On Raspberry Pi OS:

sudo apt update && sudo apt install apt-repo-sixfab
sudo apt update && sudo apt install sixfab-dx
dxrt-cli -s        # confirm device and software status
run_hello_world    # run a packaged example
dxtop              # observe the NPU in real time

From there, pick an optimised ModelZoo workload or bring your own supported ONNX model through DX-COM. The resulting .dxnn file runs through DX-RT’s C++ or Python APIs, while your Raspberry Pi application keeps handling cameras, business logic, user interfaces, networking and control.

The software stack in one view

ComponentWhat it does for you
DX-COMOptimises and compiles supported ONNX models into DEEPX .dxnn artifacts.
DX-RTRuns .dxnn models through C++ or Python APIs on the target device.
DX-StreamBuilds GStreamer-based capture, inference and post-processing pipelines.
ModelZooPre-optimised examples across detection, classification, segmentation, pose and more.
dxrt-cli / dxtopDevice status, utilisation, temperature, clocks and memory at a glance.

What can you build with it?

Start at home. A camera on your workbench that recognises when your 3D print has failed and messages you a snapshot with a one-line description. A front-door camera that tells the difference between a courier, the neighbour’s cat and someone loitering — and summarises the day’s events each evening, entirely locally. A garden or pet monitor that only alerts you about things worth seeing. All of these follow the same loop: a CNN sees the event, a compact VLM adds context, an SLM explains it or interprets your next command, and your Raspberry Pi application decides and acts.

The same loop scales to serious deployments: continuous local detection in smart cameras with reduced cloud dependence; low-power visual quality and safety monitoring beside a production line; perception and compact multimodal intelligence on robots while the Raspberry Pi runs ROS 2 and control logic; and privacy-aware occupancy and behaviour analytics in retail and smart spaces.

What it is not for

Like any edge accelerator, the AI HAT+ is not trying to compete with cloud inference. Applications that need broad world knowledge, very long conversational context or continuous learning will always run better where compute and memory are effectively unconstrained. The DX-M1M is at its best running tightly scoped, always-on intelligence next to a camera or sensor — where privacy, latency, offline operation and a single-digit-watt power budget are the requirements that actually decide the design. Being honest about that boundary is exactly what makes the platform trustworthy inside it.

Community and what’s next

Everything you need to go deeper is public: DX-AllSuite and the ModelZoo are on GitHub, and the DEEPX Developer Portal hosts documentation, guides and support channels where you can ask questions and file issues.

Where to get one

The 25 TOPS Sixfab AI HAT+ for Raspberry Pi 5 is available now from Sixfab, with the 13 TOPS version expected to launch toward the end of 2026. Documentation and the quick-start guide are linked below.

Links

No comments
Jump to the comment form

Leave a Comment