FlowEdge¶
FlowEdge is a C++23 inference runtime for robotics policies with fixed memory, deterministic execution, and explicit deadline handling. It runs two heads: flow matching (a short ODE from noise to action) and Diffusion Policy (DDIM on a Conv1D U-Net). LeRobot still owns training, cameras, and the robot driver.
- C++23 Core
- CPU · CUDA
- Tenstorrent preview
- LeRobot adapter
News¶
September 2026
-
19
Repo lives at reforcemind/FlowEdge. v0.1.2 CPU wheels are
flowedge-0.1.2-*. Not PyPI. CUDA stays a source backend. -
19
Split-K CUDA Conv1D on horizon 8 (#183). GTX 1650 matched DDIM replay stays 131 vs 345 ms vs PyTorch CUDA. Not a 10 ms loop.
-
18
Device-resident CUDA after load: flow (#167), Diffusion Policy (#168, #169), Mamba (#173). Same-process CUDA vs PyTorch CUDA (#170).
--device cudaperiod rollouts (#171). -
17
v0.1.1 CPU wheels (manylinux, Windows, macOS ARM). Not on PyPI.
--on-miss hold|drop|raise(#152). CPU fair-compare 851 vs 1409 ms,threads=1.
The problem¶
A robot loop is a deadline. At 100 Hz the motors want a new command every 10 ms. If this sample is late, they still need a defined action — hold, skip, or fault — and that choice has to be replayable (--on-miss). Matched Diffusion Policy replay is faster than PyTorch: CPU 851 vs 1409 ms, CUDA 131 vs 345 ms. That is the fair compare. The 10 ms figure is the robot’s tick, not this checkpoint’s p50.
The usual deploy path is “export the training graph and hope”: PyTorch or a general compiler, heap traffic after warmup, a dispatcher, sometimes a KV cache that grows with the horizon. That stack is the right tool for training and for broad model zoos. It is the wrong contract for a period loop with a replayable miss policy.
How FlowEdge solves it¶
Convert once. Map a pinned Hugging Face / LeRobot checkpoint into layouts the engine already knows (
backbone.*,flow.*,dp.*).Load into one arena. Weights, decode state, ODE scratch, and the thread pool are carved from a slab sized at load. Supported hot paths do not call
malloc.Sample the head. Flow matching integrates (dx/dt = v(x,t \mid c)) from noise to action (Euler / Heun / RK4). Diffusion Policy denoises an action horizon with DDIM; the RGB encoder stays outside Core. Same miss contract for both.
Honor the period.
--period-msplus--on-miss hold|drop|raise. Limits and e-stop are not in this library.
Why not the alternatives¶
Tool |
Use it for |
Not as |
|---|---|---|
LeRobot / PyTorch |
Train, encode RGB, drive the robot, compare parity |
The malloc-free period loop |
TensorRT, ONNX, ExecuTorch |
General DAGs on a given backend |
This flow or DP head’s arena, miss contract, and native kernels |
LLM / VLA servers |
Throughput, batching, GPUs |
Newest valid action before a physical deadline |
A Transformer KV cache |
Language and long context |
An open-ended robot horizon on a bounded RSS |
Flow matching is the default head because a short deterministic ODE is cheaper, in NFE, than a long DDIM walk. Diffusion Policy is first-class when that is the trained checkpoint. Mamba is the default backbone for the flow path because its state does not grow with time.
Matched PushT Diffusion Policy replay, CPU, threads=1: policy p50 851 ms vs LeRobot/PyTorch 1409 ms. Not a 10 ms loop. On GTX 1650 the same matched replay is 131 ms vs PyTorch CUDA 345 ms. Flow matching matches the same PyTorch reference on ULP (~1e-6 rel); that is not a policy p50. Performance.
Not a trainer. Not a graph runtime. Not a safety controller.
What you can use today¶
Goal |
Entry point |
|---|---|
Run flow matching |
|
Install the Python wheel |
|
Deploy LeRobot Diffusion Policy |
|
Numbers vs PyTorch |
|
Convert a checkpoint |
|
Keep your encoder, run only the head |
|
Optional local IPC + deadlines |
|
Cached SmolVLA expert (VLM in LeRobot) |