Updated October 7, 2026. This guide replaces our 2024 Raspberry Pi AI Kit walkthrough. The new edition draws on current vendor documentation, public code and other developers’ experiments; we have not rerun these examples or benchmarks on hardware for this update.

A camera drawing boxes around people is a useful first test. The next questions are harder: can you count a person once, record an animal before it leaves the frame, or keep several cameras running without filling a queue? That is where an accelerator’s software examples become more useful than its TOPS rating.

The examples below cover detection, tracking, recording and custom models on the original Hailo-8L kit and newer boards. The AI HAT+ 2 also opens a path to local language and vision-language applications.

Raspberry Pi AI Kit, AI HAT+ or AI HAT+ 2?

The original AI Kit pairs an M.2 HAT+ with a Hailo-8L module. Raspberry Pi’s current hardware documentation marks that kit as out of production and points new designs toward the integrated AI HAT boards.

Scroll horizontally to explore all columns →

Research data
BoardAcceleratorWhat it is useful for

Original Raspberry Pi AI Kit

Hailo-8L · 13 TOPS, INT8

Keep using an existing kit for compatible computer vision examples.

AI HAT+ · 13 TOPS

Hailo-8L · INT8

A starting point for supported detection, pose and segmentation models.

AI HAT+ · 26 TOPS

Hailo-8 · INT8

More vision inference capacity, where the model and pipeline can use it.

AI HAT+ 2

Hailo-10H · 40 TOPS INT4 / 20 TOPS INT8 · 8 GB onboard memory

Vision plus supported local LLMs and VLMs. Its memory belongs to the accelerator.

The Hailo-10H product brief specifies up to 40 TOPS at INT4 and 20 TOPS at INT8. Precision matters when comparing those figures with the INT8 ratings of the 8 and 8L. Raspberry Pi describes the AI HAT+ 2’s vision performance as comparable to the 26-TOPS AI HAT+; individual models and workloads can behave differently. Choose a supported application first, then look at measurements of that workload.

Our recommendation is to keep a working 8L or 8 setup for a supported CV workload. Consider the 10H when your project specifically needs its GenAI capabilities, or when a benchmark of your chosen vision model justifies the change.

Choose a Hailo example by task

This directory combines small official demos, application integrations and community projects. “8L / 8 / 10H” identifies the applicable hardware family; an individual model must still have a matching compiled file. The walkthroughs below explain where each example fits and what to check before adapting it.

Scroll horizontally to explore all columns →

Research data
TaskStart hereHardware / scope

First camera detection

rpicam-apps and hailo-detect-simple

8L / 8 / 10H · official demos

Tracking and counting

hailo-detect and a counting callback

8L / 8 / 10H · application logic required

Pose and segmentation

hailo-pose and hailo-seg

8L / 8 / 10H · model-dependent

Python camera control

Picamera2 detection example

Pi camera + matching HEF

Small objects and several cameras

Tiling and multisource examples

8L / 8 / 10H · check available models

Camera events and recording

Frigate

Documented detector: 8L / 8

Wildlife video capture

Pi_Hailo_Wildlife_3

Author documents 8L / 10H

Train for your own objects

RasPi_YOLO and Hailo’s barcode example

Compile for the exact target chip

Newer YOLO experiments

YOLO26 with conversion and evaluation code

Community benchmark: 8L

Local text and image questions

Simple LLM Chat and VLM Chat

10H · GenAI dependencies

Hardware setup: establish a clean baseline

Use a Raspberry Pi 5, suitable power supply, storage, cooling and one compatible accelerator board. A CSI camera such as Camera Module 3 is useful for the camera walkthrough; file and USB-camera examples offer other input routes.

Disconnect power before assembly. Follow Raspberry Pi’s illustrated HAT mounting instructions for the ribbon cable, spacers and cooling. Check case clearance before buying an enclosure. If an NVMe adapter already uses the Pi’s PCIe connector, plan the expansion arrangement before adding the accelerator.

For an existing deployment, use a spare system image to try the current stack. Record the working OS and package versions before upgrading; it is much easier to compare two known setups than to reconstruct a half-updated one.

Software setup: install the stack for your chip

The current Raspberry Pi AI installation guide uses 64-bit Raspberry Pi OS Trixie. Its two package paths are mutually exclusive: hailo-all for AI Kit / AI HAT+, and hailo-h10-all for AI HAT+ 2.

Update the OS and firmware, then reboot:

Bash
sudo apt update
sudo apt full-upgrade -y
sudo rpi-eeprom-update -a
sudo reboot

Choose one installation block.

Hailo-8L or Hailo-8:

Bash
sudo apt install dkms hailo-all
sudo reboot

Hailo-10H / AI HAT+ 2:

Bash
sudo apt install dkms hailo-h10-all
sudo reboot

Confirm that the runtime identifies the device:

Bash
hailortcli fw-control identify

Check the architecture reported against your board. Raspberry Pi recommends enabling PCIe Gen 3 for the original AI Kit; AI HAT+ and AI HAT+ 2 apply that setting automatically. Follow the configuration steps in the installation guide.

Where did hailo-rpi5-examples go?

The hailo-rpi5-examples repository now labels itself outdated and directs readers to hailo-apps. If you arrived looking for basic_pipelines, detection.py or detection_counts.py, keep the task you want to solve, but start from the current application structure.

After the device check passes, Hailo’s shared installation path for Raspberry Pi is:

Bash
git clone https://github.com/hailo-ai/hailo-apps.git
cd hailo-apps
./install.sh
source setup_env.sh

The installer can request elevated privileges. Read its prompts and review the repository before running it. Activate the environment in each new terminal. For reproducibility, record the commit and runtime:

Bash
git rev-parse HEAD
hailortcli --version

We reviewed hailo-apps at commit c61f843, dated October 5, 2026. Its 26.10.0 changelog includes installation, resource compatibility and camera fixes. A newer repository checkout and an older OS runtime are not automatically a compatible pair; consult the application’s prerequisites when combining versions.

Hailo RPi5 examples: detection, tracking, pose and segmentation

1. Object detection: start with one input

Install rpicam-apps with sudo apt install rpicam-apps, then check the camera without inference using rpicam-hello. Raspberry Pi’s packaged YOLOv8 camera demo adds detection with:

Bash
rpicam-hello -t 0 \
  --post-process-file /usr/share/rpi-camera-assets/hailo_yolov8_inference.json

For an application you can extend, Hailo’s simple detection pipeline provides inference, a callback and display without the full example’s tracker. Run it from the activated Hailo Apps environment:

Bash
hailo-detect-simple --input rpi --show-fps

The shared application reference also supports inputs such as usb, an RTSP URL or a local video path. A saved clip is a useful first input because every code change can be checked against the same frames.

2. Object tracking and people counting

The full detection example adds tracking and Python callbacks. Start with hailo-detect --input rpi --show-fps, then inspect the callback and tracked metadata before adding business logic.

A per-frame count answers “how many people are visible now?” A doorway counter needs track IDs, a crossing line, a direction and a rule for when to count. Counting every person-shaped box on every frame would count the same person repeatedly. Test crossings, reversals and temporary occlusion separately; a tracker does not make those decisions for you.

3. Pose estimation

Hailo’s pose example is a starting point for keypoint overlays and movement interfaces:

Bash
hailo-pose --input rpi --show-fps

A skeleton overlay is only the first layer of a gesture or exercise application. Define the movement, camera angle and failure cases before deriving a score from the keypoints.

4. Instance segmentation

Use the segmentation example when you need the visible shape of an object rather than a rectangular box:

Bash
hailo-seg --input rpi --show-fps

Mask decoding and drawing also consume host resources. Compare the complete application with and without its overlays before concluding that inference is the bottleneck.

Camera Module 3 and Python with Picamera2

5. A small Python camera application

Raspberry Pi’s Picamera2 Hailo detection example captures a lower-resolution inference stream, calls Hailo.run() and draws results on the preview. It is a small entry point for developers who want to own the camera loop. The examples have moved to the separate picamera2-examples repository, and the current file selects different default model paths for Hailo-10H and earlier hardware.

Read this example alongside its coco.txt label file and the models installed on your Pi. It exposes the handoff between camera frames, inference output and display coordinates more directly than a large application does.

For a custom camera pipeline, check focus, exposure, resizing and channel order with a known image. A fast model fed the wrong input format can produce convincing-looking but incorrect detections. Keep one saved frame and its expected objects as a debugging fixture.

Small objects, multiple cameras and custom tracking

6. Tiling for small objects

The Hailo tiling example processes crops of a larger image and combines the detections. Its default demonstration uses aerial footage. Start with the supplied example before changing camera input:

Bash
hailo-tiling
# A separate run with additional scales:
hailo-tiling --multi-scale

Tiling can preserve detail that resizing a whole frame would discard, but each extra tile adds inference work. Inspect objects near tile boundaries and measure latency as you increase the grid.

7. Several camera streams

The multisource pipeline demonstrates per-stream decoding, shared accelerator work and routing results back to the right output. Its input arguments differ from the single-camera apps; begin with hailo-multisource --help and the source examples in its README.

The current guide requires sources of the same type and suggests starting with lower frame rates on Raspberry Pi. Measure each stream’s delay and dropped frames. An aggregate FPS number can hide one camera falling behind.

Our original detection-tracking example

The 2024 edition introduced DRL’s custom detection-tracking pipeline, combining segmentation with ORB features and Lucas–Kanade optical flow to bridge missed detections. We retain the code as a historical reference. Its old submodule and environment instructions have not been revalidated against the 2026 stack.

For a new implementation, first try the current official tracker. If you need optical-flow fallback, define when a track expires: carrying a box forward indefinitely can turn one detection into a persistent false result.

Community projects worth opening

8. Frigate: turn detections into camera events

Frigate’s Hailo integration is a practical route from a demo to an application with video events. Its documented detector supports Hailo-8 and Hailo-8L, can select a default model for the detected architecture, and caches it after download.

Use the hardware installation and detector configuration for your Frigate release; the documented scope here is 8/8L. Its detector timing covers one part of a larger workload that also includes video decoding and recording.

9. Wildlife capture with a pre-recording buffer

Gordon999’s Pi_Hailo_Wildlife_3 combines Picamera2 detection and circular recording to save MP4 clips, including roughly five seconds before a trigger. The author documents Hailo-8L and Hailo-10H setups, configurable object classes and a camera-specific configuration.

One detail makes the project particularly instructive: its author notes that a red squirrel can be labelled as a bear or cat because the supplied classes do not include it. Recording a useful event and correctly identifying a species are separate objectives. Adapt the camera settings and dependency installation to your own environment instead of copying system-wide package overrides.

10. Custom YOLO: from a dataset to a HEF

Luke Ditria’s RasPi_YOLO follows dataset preparation, training, ONNX export, calibration and deployment. Its Hailo route explicitly selects hailo8 or hailo8l; its Sony IMX500 route is a different deployment target.

Use it to understand the conversion workflow, then reconcile its software versions with your current toolchain. For another worked example, Hailo’s barcode retraining guide includes training and compilation notebooks. Detecting the location of a barcode is still different from decoding its contents.

11. YOLO26: inspect the evaluation as well as the demo

Daniel Dubinsky’s YOLO26 on Hailo-8L repository publishes conversion scripts, Python/C++ inference and COCO evaluation code. It is useful for studying the conversion and evaluation workflow. Check the timing boundaries before reusing its FPS figures: the C++ benchmark prepares images before its timed loop, so a quoted rate needs to be tied to the specific script and run.

Hailo also provides an official YOLO26 example that splits neural inference into a HEF and post-processing into ONNX Runtime. This is an advanced route: inspect the current files and model resources before adapting it. A model family name alone does not tell you where all computation runs.

AI HAT+ 2 examples: local LLMs and visual questions

The following examples use the Hailo-10H GenAI path. They are not upgrades you can enable on an 8L by installing another Python package.

12. Start with a small text request

The Simple LLM Chat example exposes a minimal request and response. After the shared installation, activate its environment and install the optional GenAI dependencies:

Bash
source setup_env.sh
python -m pip install -e ".[gen-ai]"
python -m hailo_apps.python.gen_ai_apps.simple_llm_chat.simple_llm_chat --list-models
python -m hailo_apps.python.gen_ai_apps.simple_llm_chat.simple_llm_chat

Model downloads can occur on first use. Select from the supported resources for the application and chip; the presence of a model on Hugging Face does not establish that it will run on this accelerator.

13. Ask a question about a camera frame

The VLM Chat application accepts a live camera and lets you ask about a captured frame:

Bash
python -m hailo_apps.python.gen_ai_apps.vlm_chat.vlm_chat --input rpi

Its interface keeps camera handling separate from answering a question. A responsive video window does not imply that a new language answer is generated for every frame. Measure time to the first answer and verify factual errors on your own scenes.

Once those components work independently, the voice assistant example offers a more involved speech-to-text, LLM and speech-output application. Treat it as a composition exercise: inspect which stages run on the accelerator and which still use the host.

Hailo benchmarks: official figures and Raspberry Pi tests

Official Model Zoo: YOLOv8s throughput

Hailo publishes the following FPS values for its standard yolov8s entry with a 640×640×3 input. All three source tables name an Intel Core i5-9400 host at 2.90 GHz. These are vendor reference results, not Raspberry Pi camera measurements. The links below pin the Model Zoo release; the compiled artifacts differ across chip families.

Scroll horizontally to explore all columns →

Research data
Chip and published sourceFPS · batch 1FPS · batch 8

Hailo-8L · Model Zoo v2.19.1

110

208

Hailo-8 · Model Zoo v2.19.1

491

491

Hailo-10H · Model Zoo v5.4.0

166

252

Batching can increase throughput without improving the response time of a single frame. The table also illustrates why TOPS alone is a poor model selector: the 40-TOPS headline for the 10H does not predict its position for this particular INT8 vision workload.

A Raspberry Pi 5 reference point

In a June 2024 Hailo community post, user omria reported 127.85 FPS for YOLOv8s at 640×640 on a Pi 5 with an 8L, PCIe Gen 3 and batch size 8. The follow-up identifies hailortcli run {hef} --batch-size 8 as the command. This is a historical model-throughput result; the post does not specify a complete software/HEF revision or measure a live camera application.

Language models: first response versus a long answer

In a Hailo-authored article published by Raspberry Pi on February 25, 2026, the vendor reports time to first token of 320 ms on Hailo-10H versus 2,039 ms on the Pi 5 CPU for Qwen2.5-1.5B at 4-bit precision with 96 prefill tokens. The CPU used llama.cpp. This supports a specific responsiveness claim, not a general token-generation speedup; the post does not give a complete reproducibility package.

For a concrete example of request-level timing, CNX Software’s January 20 review includes a DeepSeek-R1-Distill-Qwen-1.5B response on a Pi 5 with 2 GB RAM and AI HAT+ 2. Its output records 380 tokens over 56.508 seconds of total request time: approximately 6.72 tokens/s. That single-request rate includes more than decoding. The review’s CPU comparison uses a different timing denominator and a CM5 setup, so it cannot establish a controlled hardware speed ratio.

These measurements answer different application questions. A short command benefits from a quick first response; a long answer also needs adequate sustained generation speed. The 10H’s dedicated memory and ability to offload work from the host may matter even when a particular model’s generation rate is modest.

A useful test for your own project

Use the same clip, model variant, resolution and accuracy thresholds for every run. Save the HEF, repository revision and runtime versions. Then measure:

  • CV: completed frames, dropped frames, camera-to-result delay and accuracy on representative images.
  • LLMs and VLMs: cold start, time to first token, generation rate, answer length and task correctness.
  • The system: CPU load, memory, temperature and wall power over a sustained run.

For example, a 30 FPS display may be camera-limited, while a 100 FPS model benchmark may exclude decoding. Neither result alone tells you whether a two-camera application meets a 100 ms deadline.

Custom models, Python and common setup failures

The old edition’s “Python support is coming” guidance no longer applies. Current Picamera2 and Hailo Apps examples expose Python inference. What still needs care is the contract between the model, runtime and application.

The Hailo Model Zoo compatibility notice directs Hailo-8/8L users to its 2.x line with Dataflow Compiler 3.x; the master line targets Hailo-10/15. Keep those compiler/model branches separate from the version of the applications repository.

A typical custom-model path is training → export → calibration and quantization → compilation for the target chip → accuracy evaluation → application integration. A successful compile does not establish accuracy. Check label mapping, tensor shapes, preprocessing and the post-processing expected by your application before replacing its default HEF.

Scroll horizontally to explore all columns →

Research data
SymptomFirst check

Device is not detected

Power down before checking the PCIe ribbon. Confirm the correct package family, reboot and inspect the device-identification output and kernel logs.

Runtime / driver mismatch

Compare installed driver and HailoRT versions; inspect old package holds and mixed installation sources. Use the documented compatible set.

HEF is rejected or a model is missing

Check the target architecture, runtime compatibility and resource download. A filename containing “YOLO” is insufficient.

Python cannot import a Hailo module

Confirm which virtual environment is active. Reconcile PyHailoRT and TAPPAS bindings with the runtime; avoid unrelated pip packages with similar names.

Camera works but detections look wrong

Check channel order, resizing, normalization, labels and post-processing with a known image.

FPS is low or the display lags

Separate decode, inference, callbacks and rendering; inspect queues. Check temperature and input frame rate before changing the model.

Hailo versus Coral or Jetson

Start the comparison with the deployment format. Coral’s Edge TPU requires compatible quantized TensorFlow Lite models, compiled for that accelerator; Hailo uses compiled HEFs. Account for the model conversion and application code you would need to change when moving between them.

A Jetson Orin Nano developer kit provides a complete NVIDIA development platform. Compare its hardware and software dependencies with the complete Pi setup, including the camera, storage, cooling and accelerator. A board’s headline throughput cannot answer whether the application you want is already supported.

Hailo computer vision FAQ

These answers cover setup, custom models and camera pipelines on Raspberry Pi and Hailo. They draw on recurring community problems and current documentation, checked in October 2026. Commands belong to the specific examples linked below; we have not reproduced these setups on hardware.

Why is my Raspberry Pi 5 not detecting the Hailo accelerator?

Start with hailortcli fw-control identify, before running a model. If it fails, check the PCIe connection, power and driver installation using Raspberry Pi’s AI setup guide. Disconnect power before reseating the ribbon cable. Use the package family for your board: hailo-all for Hailo-8/8L, or hailo-h10-all for Hailo-10H. Read the accompanying error: a missing device, a driver mismatch and an occupied accelerator call for different checks.

How do I fix a Hailo driver and library version mismatch?

The loaded kernel driver and the HailoRT library used by your application must be compatible. A reported mismatch after a manual upgrade illustrates how mixing Raspberry Pi packages with separately downloaded releases can break that pairing. Follow one supported installation route, including its Python bindings, then reboot after a driver update. Compare the versions named in the error with the installation instructions for your board; copying a version number from an old forum fix can recreate the mismatch.

Why does Python report “No module named hailo” or “hailo_platform”?

Check which Python interpreter is running the script. For a shared hailo-apps installation, run source setup_env.sh from the repository root in each new terminal, as described in Hailo’s environment setup instructions. hailo_platform provides the HailoRT Python API; pipeline examples also use the separate hailo metadata bindings. A working command-line device check does not establish that both are available inside your virtual environment. Install the components required by that example and runtime version.

Where are hailo-rpi5-examples and detection_counts.py now?

The old hailo-rpi5-examples repository is marked outdated and points to hailo-apps. Old filenames do not necessarily have a direct replacement. For detection and tracking, start with the current detection callback, which exposes labels, confidence and tracking IDs. If your goal is a cumulative count, add counting logic to that callback; the number of detections in one frame is not the number of visitors over time.

Why does my HEF fail with HAILO_INVALID_HEF?

Inspect the file with hailortcli parse-hef model.hef and compare its target with hailortcli fw-control identify. A Hailo-8 HEF loaded on Hailo-8L can produce this error. Also check compiler/runtime compatibility and the file itself: another resolved invalid-HEF report involved a HAR archive saved with a .hef extension. Renaming an archive does not convert it into an executable model. Obtain or compile a HEF for the intended device.

Can I compile a HEF on the Raspberry Pi itself?

The documented Dataflow Compiler workflow runs on a supported Linux x86_64 build machine. The Pi runs the resulting HEF through HailoRT. Hailo’s compiler-versus-runtime explanation makes that distinction; the current Ultralytics Hailo export guide retains the same architecture requirement. Compile on a compatible workstation, then transfer the HEF and any application metadata to the Pi. You do not need the Dataflow Compiler installed on the Pi for inference.

How do I convert a custom YOLO model to HEF?

For supported Ultralytics architectures, the current direct Hailo exporter accepts custom .pt weights. On a supported build machine with the matching Dataflow Compiler installed, a Hailo-8L export can use:

Bash
yolo export model=best.pt format=hailo name=hailo8l data=dataset.yaml

Replace the files with your weights and representative calibration dataset. Check the exporter’s model-support table and select your actual chip. An arbitrary ONNX model still needs a compatible parsing, optimization and compilation workflow; changing its extension to .hef is not conversion.

Why does my compiled model run but detect nothing?

A successful inference call does not prove that the application decodes the output correctly. Check input resizing, color order and normalization, then the output decoder, class count, labels and confidence threshold. Hailo’s custom YOLO retraining walkthrough includes model-specific postprocessing configuration. Compare the original model and HEF on the same known images. If accuracy dropped during quantization, review whether calibration images represent your cameras and objects before changing thresholds to hide the loss.

What causes “Input buffer size is different than expected”?

Inspect the input shape and format reported by the loaded model, then the actual array’s shape, dtype, byte count and memory layout. Do not assume the original training tensor layout is the runtime input layout. The current Hailo inference helper shows how it reads HEF input information and creates bindings. If Python reports zero bytes despite a correctly sized array, reduce the problem to that binding step and check the supported runtime, Python-wheel and NumPy combination before changing the model.

Can I use OpenCV and Python instead of GStreamer?

Yes. Hailo provides a standalone Python object-detection example that uses HailoRT and OpenCV without GStreamer. For a Raspberry Pi camera, the Picamera2 Hailo example is another starting point. Keep input preprocessing and output decoding consistent with the chosen HEF. A GStreamer callback is a different interface: hailo.get_roi_from_buffer() reads metadata from a GStreamer buffer, not directly from an OpenCV NumPy image.

Why is live video delayed even when inference FPS is high?

Model throughput and the age of a displayed frame measure different things. Capture, decoding, Python callbacks or rendering can leave frames waiting in queues. Timestamp frames at capture and at the result to locate the delay. For an application that needs the freshest frame, GStreamer’s queue limits and leaky modes let you bound buffering and discard older queued buffers. That tradeoff drops frames, so validate it against tracking and counting requirements. A high inference FPS alone does not establish low end-to-end latency.

Can I run Hailo detection over SSH without a monitor?

Yes, with an application that supports headless output. The current standalone object-detection app accepts --no-display. After completing its installation and activating the environment, run this from the hailo-apps repository root, replacing the HEF path:

Bash
python hailo_apps/python/standalone_apps/object_detection/object_detection.py \
  -n /path/to/model.hef -i usb --no-display

This example requires a compatible detection HEF with HailoRT postprocessing. The flag belongs to this app; do not assume every GStreamer demo accepts it.

Can one Hailo accelerator process multiple cameras?

Yes. Hailo’s multisource application decodes and scales streams, feeds frames through the accelerator and routes results back to the corresponding source. Its current example requires sources of the same type and uses --sources; check hailo-multisource --help for the installed version. Measure the whole pipeline with your camera resolution and rate. Adding streams shares accelerator and host resources, so a single-camera FPS result is not a per-camera guarantee.

Why do I get HAILO_OUT_OF_PHYSICAL_DEVICES when the board is connected?

The error can mean that no device is free for the requested allocation. One documented cause is an application that still holds the device. Identify and stop that application cleanly, and make sure your own code releases its HailoRT resources. If nothing holds the device, return to driver and device-visibility checks. For deliberate concurrent workloads, study the shared VDevice and scheduler configuration instead of opening competing exclusive sessions.

Why does Hailo work on the host but fail inside Frigate Docker?

Check both the host driver and the container’s access to the accelerator. Frigate’s Hailo installation instructions require passing /dev/hailo0 into Docker and include Raspberry Pi OS-specific driver steps. Follow the requirements for the Frigate release you run, including compatible HailoRT versions. Installing packages inside a container does not replace the host’s loaded kernel driver. Once the device is accessible, verify the Hailo detector and model configuration in Frigate.

How do I count people or vehicles without counting them again in every frame?

Use tracking plus an explicit counting rule. The Hailo detection callback exposes a tracking ID through HAILO_UNIQUE_ID. For a doorway or road, count a track when it crosses a defined line in the intended direction, and remember that event to avoid counting it on the next frame. Test occlusion, re-entry and lost tracks with representative footage. Tracking IDs can change; they are not permanent identities. Report counting accuracy separately from detector FPS.

← Back to blog