Baselines¶
netforge_rl.baselines.
Policy |
|
|---|---|
|
uniform |
|
isolate first compromised host, else analyse |
|
exploit without recon — usually 0 compromises |
|
recon → exploit → pivot; ~2–3 hosts/ep, SLA drops |
|
on-device IPPO ( |
ExploitRemoteService needs prior DiscoverNetworkServices. Use kill-chain as
the Red reference, not heuristic-red.
python -m benchmarks.build_leaderboard --episodes 5 --max-steps 150
python -m benchmarks.train_curve --name blue_ransomware --iters 40
Committed IPPO on ransomware: mean reward 0.06 → 0.71 (40 iters, CPU).
Load with jax_ppo.load_params. RLlib: benchmarks/rllib_rmappo.py.
Run · Self-play / Elo