NetForge

Cybersecurity gym for reinforcement learning. Red compromises hosts. Blue contains them from SIEM, not from an oracle. Train with PettingZoo or Gymnasium.

  • Red vs Blue
  • SIEM telemetry
  • PettingZoo
  • Gymnasium
1 Red · 3 Blue DMZ, internal, restricted
SIEM only Blue never sees the true map
durative actions exploits and isolates take ticks

News

September 2026

  • Gym ids: PettingZoo netforge/ransomware-v4, Gymnasium NetForge/Blue-v4.

  • netforge CLI: run, evaluate, questions, HTML replay.

  • Arena train / dev / hidden. Typed EnvConfig. Changelog

The problem

The job is a SOC shift, not a capture-the-flag. Blue only ever sees a log buffer: Sysmon-like events that can arrive late, drop, or never fire if Red stays quiet. Isolation is slow and takes hosts off the mission. A policy that reads the simulator’s compromise map, or that wins by unplugging the plant, is not a defender.

Most cyber RL gyms still train a different job. Blue gets something close to true host state. Actions finish in one step. The paper reports mean reward against one scripted Red. The agent looks strong there and fails as soon as the logs are incomplete or the attacker changes.

NetForge is a gym for the first job: train and rank on SIEM, SLA, and a Red population. Trainers stay in your repo.

Compared to other gyms

CybORG / CAGE is the competition stack papers already cite. CyberBattleSim and NASim are graph capture and pentest. Yawning Titan is abstract graph defense. NetForge is the SOC gym: delayed logs, durative actions, PettingZoo ids, Arena metrics.

Blue observation

Time

What you rank

API

CybORG / CAGE

Host table, often close to true state

Mixed; many actions resolve in-step

Historically mean return vs scripted Red (B-line, Meander)

Custom env + wrappers

CyberBattleSim

Discovered attack graph

Instant node/credential actions

Red capture; Blue is thin

Gymnasium

NASim

Scan-revealed network

Instant exploits

Attacker success

Gymnasium

Yawning Titan

Abstract graph nodes

Instant

Graph defense

Gymnasium

NetForge

SIEM buffer + belief graph. Oracle is diagnostic only

Durative: exploits and isolates take ticks

Arena: mission, SLA, FPs, security, CVaR vs a Red population

PettingZoo + Gymnasium

NetForge does not try to beat CybORG on host-type count. LICENSE still carries CybORG / DSTG notices. The bet is the observation and the score: Blue trains on logs, isolate-all wrecks SLA, and one campaign is not an eval.

Ranges, Caldera, and FARLAND are emulation or adversary tooling. Use them when you need packets on a wire. This repo is the pip-installable gym.

What NetForge does

Red starts blind, discovers hosts, exploits CVEs, escalates, exfiltrates or hits a PLC. Blue isolates, restores, drops decoys, and reads Sysmon-like logs that can be late or noisy. Padding hosts (169.254.x.x) are decoys.

Red, Blue, and the network reset, observe SIEM, act, next tick God-mode vs SIEM, instant logs vs delay, isolate-all vs SLA

Install

python -m pip install 'netforge-rl @ git+https://github.com/reforcemind/NetForge_RL'
netforge run hospital_ransomware --replay replay.html
import netforge_rl
import gymnasium as gym
from pettingzoo import make

env = make('parallel', 'netforge/ransomware-v4', max_ticks=80)
obs, infos = env.reset(seed=0)
blue = gym.make('NetForge/Blue-v4', max_ticks=80)

Not on PyPI. Python 3.12+. Trainers stay in your repo.

MIT, with CybORG / DSTG notices in the LICENSE file.