FlowEdge

FlowEdge

FlowEdge is a C++23 inference runtime for robotics policies with fixed memory, deterministic execution, and explicit deadline handling. It runs two heads: flow matching (a short ODE from noise to action) and Diffusion Policy (DDIM on a Conv1D U-Net). LeRobot still owns training, cameras, and the robot driver.

  • C++23 Core
  • CPU · CUDA
  • Tenstorrent preview
  • LeRobot adapter
851 vs 1409 CPU DDIM p50 · ms · threads=1
131 vs 345 CUDA DDIM p50 · ms · GTX 1650
malloc = 0 Supported hot paths after load

News

September 2026

  • 19

    Repo lives at reforcemind/FlowEdge. v0.1.2 CPU wheels are flowedge-0.1.2-*. Not PyPI. CUDA stays a source backend.

  • 19

    Split-K CUDA Conv1D on horizon 8 (#183). GTX 1650 matched DDIM replay stays 131 vs 345 ms vs PyTorch CUDA. Not a 10 ms loop.

  • 18

    Device-resident CUDA after load: flow (#167), Diffusion Policy (#168, #169), Mamba (#173). Same-process CUDA vs PyTorch CUDA (#170). --device cuda period rollouts (#171).

  • 17

    v0.1.1 CPU wheels (manylinux, Windows, macOS ARM). Not on PyPI. --on-miss hold|drop|raise (#152). CPU fair-compare 851 vs 1409 ms, threads=1.

The problem

A robot loop is a deadline. At 100 Hz the motors want a new command every 10 ms. If this sample is late, they still need a defined action — hold, skip, or fault — and that choice has to be replayable (--on-miss). Matched Diffusion Policy replay is faster than PyTorch: CPU 851 vs 1409 ms, CUDA 131 vs 345 ms. That is the fair compare. The 10 ms figure is the robot’s tick, not this checkpoint’s p50.

The usual deploy path is “export the training graph and hope”: PyTorch or a general compiler, heap traffic after warmup, a dispatcher, sometimes a KV cache that grows with the horizon. That stack is the right tool for training and for broad model zoos. It is the wrong contract for a period loop with a replayable miss policy.

Matched replay loading, CUDA then CPU. FlowEdge CUDA 131 ms vs PyTorch 345 ms; FlowEdge CPU 851 ms vs PyTorch 1409 ms.

How FlowEdge solves it

  1. Convert once. Map a pinned Hugging Face / LeRobot checkpoint into layouts the engine already knows (backbone.*, flow.*, dp.*).

  2. Load into one arena. Weights, decode state, ODE scratch, and the thread pool are carved from a slab sized at load. Supported hot paths do not call malloc.

  3. Sample the head. Flow matching integrates (dx/dt = v(x,t \mid c)) from noise to action (Euler / Heun / RK4). Diffusion Policy denoises an action horizon with DDIM; the RGB encoder stays outside Core. Same miss contract for both.

  4. Honor the period. --period-ms plus --on-miss hold|drop|raise. Limits and e-stop are not in this library.

Hardware runs the loop, FlowEdge runs the policy, LeRobot trains and drives the robot convert, load arena, sample flow or DDIM, period, act

Architecture · Relay · Flow matching vs DDIM

Why not the alternatives

Instead of PyTorch-in-the-loop, TensorRT/ONNX, a growing KV cache, long DDIM, or model servers — FlowEdge runs this flow or DP head in a sized arena with a miss contract

Tool

Use it for

Not as

LeRobot / PyTorch

Train, encode RGB, drive the robot, compare parity

The malloc-free period loop

TensorRT, ONNX, ExecuTorch

General DAGs on a given backend

This flow or DP head’s arena, miss contract, and native kernels

LLM / VLA servers

Throughput, batching, GPUs

Newest valid action before a physical deadline

A Transformer KV cache

Language and long context

An open-ended robot horizon on a bounded RSS

Flow matching is the default head because a short deterministic ODE is cheaper, in NFE, than a long DDIM walk. Diffusion Policy is first-class when that is the trained checkpoint. Mamba is the default backbone for the flow path because its state does not grow with time.

Matched PushT Diffusion Policy replay, CPU, threads=1: policy p50 851 ms vs LeRobot/PyTorch 1409 ms. Not a 10 ms loop. On GTX 1650 the same matched replay is 131 ms vs PyTorch CUDA 345 ms. Flow matching matches the same PyTorch reference on ULP (~1e-6 rel); that is not a policy p50. Performance.

Not a trainer. Not a graph runtime. Not a safety controller.

What you can use today

Goal

Entry point

Run flow matching

Getting Started

Install the Python wheel

Python

Deploy LeRobot Diffusion Policy

LeRobot adapter

Numbers vs PyTorch

Performance

Convert a checkpoint

Converter

Keep your encoder, run only the head

Capabilities

Optional local IPC + deadlines

Relay quickstart

Cached SmolVLA expert (VLM in LeRobot)

Transformer / SmolVLA