Benchmark map¶
Suite |
Location |
Measures |
|---|---|---|
Kernels |
|
Dense ops, including DP Conv1D shapes |
Runtime |
|
Relay admission, queues, delivery |
Flow matching vs PyTorch |
|
ULP / rel-error on |
Diffusion Policy vs PyTorch |
|
Same observations, processor, noise, DDIM vs LeRobot |
Period loop |
|
|
Closed loop |
|
PushT outcome — not a timing proof |
SmolVLA expert |
|
Cached-VLM expert vs source; VLM stays in PyTorch |
Transformer fixture |
|
GPT-2 smoke vs Hugging Face; host decoder p50 is not a policy result |
ONNX companion |
|
Fixed-shape ORT adapter; measure separately |
The two product heads are flow matching and Diffusion Policy. Without FlowEdge you keep them in PyTorch/LeRobot or a graph compiler. With FlowEdge you convert the head, pin RSS, and compare against that same PyTorch reference. Mix ULP with p50 and the claim is wrong. Performance.
Canonical policy report¶
python -m flowedge_dev bench policy models/policy.safetensors \
--source models/diffusion_pusht \
--revision <immutable-hf-revision> \
--steps 10 --iterations 100 --warmup 5 --threads 1 \
--output bench/artifacts/policy/diffusion-pusht-report.json
JSON + Markdown together. Reference is LeRobot/PyTorch; candidate is FlowEdge native.
Interpretation¶
Claim |
Allowed when |
|---|---|
PyTorch speedup |
Same checkpoint, processor, host, threads, schedule |
Real-time |
Full period, delivery path, miss counts |
Zero hot allocations |
Native path instrumentation |
Hardware result |
Artifact names the board |
CPU available · CUDA matched PushT replay on GTX 1650 · Tenstorrent planned. CUDA policy p50 is not the README CPU headline. Not TensorRT/ONNX. Not Jetson/ARM (#70).