English | 简体中文
BVR-Sim is a 3D beyond-visual-range air-combat environment for reinforcement-learning research, with Python and accelerated C++ backends and a built-in skrl PPO training pipeline for immediate experimentation.
| Heterogeneous F-22/F-16 6v6 | Human–AI Real-Time Air Combat | Heterogeneous Air Combat + Ground Air Defense |
|---|---|---|
| Well-Tuned IPPO: 2v2 | Composite Missile Guidance + FDM | UnrealCV + UE5 Real-Time Rendering |
|---|---|---|
The shared tactical interface maps heading, altitude, speed, and fire commands to aircraft-specific inner-loop controllers. Simulation, learning interfaces, and optional telemetry or rendering remain decoupled, allowing the same scenario to run through the Python reference backend or accelerated C++ backend.
The C++ JSBSim backend currently provides mappings for the following aircraft families:
- F-15
- F-16
- F/A-18
- F-22
- F-4N Phantom II
- AJ 37 and JA 37 Viggen
Different aircraft models, controller parameters, sensors, and weapon loadouts can coexist in the same scenario.
The following comparison reproduces Table 1 of the paper for the cited public releases. A cross means that documented or direct support was not found in the corresponding release.
| Capability | BVR Gym [3] | LAG [4, 5] | B-ACE [6] | BVR Sim |
|---|---|---|---|---|
| Open source | ✓ | ✓ | ✓ | ✓ |
| JSBSim aircraft dynamics | ✓ | ✓ | × | ✓ |
| BVR-class missile engagement | ✓ | × | ✓ | ✓ |
| Mixed aircraft models in one scenario | × | × | × | ✓ |
| Built-in maneuver-and-fire interface | Maneuver only | ✓ | ✓ | ✓ |
| Interchangeable Python/native C++ backends | × | × | × | ✓ |
| Uniform fixed-width entity table | × | × | × | ✓ |
| Rule baseline and self-play | ✓ | ✓ | ✓ | ✓ |
| ACMI/Tacview export or telemetry | ✓ | ✓ | × | ✓ |
| Real-time 3D visualization | FlightGear | Tacview | Godot | Native viewers |
The archived HAPPO [8] and MAPPO [7] traces above reproduce Fig. 3 of the paper and verify end-to-end integration on the same 2-vs-2 BVR task. Each curve is from one run, so the figure is an integration check rather than an algorithm ranking.
BVR Sim provides two complementary flight-dynamics paths for different stages of reinforcement-learning development.
The JSBSim FDM path provides higher-fidelity, near-realistic aircraft dynamics for flight-control and engagement studies. Simple FDM is a lightweight alternative for fast RL pipeline validation, reward debugging, and early convergence checks.
Table 3 of the paper reports headless throughput at a 0.4-second decision interval. Values are mean ± standard deviation over three repeats.
| Scale | Python steps/s | C++ steps/s | Speedup |
|---|---|---|---|
| 1v1 | 95.84 ± 8.07 | 260.65 ± 2.75 | 2.72× |
| 2v2 | 51.01 ± 8.09 | 143.58 ± 4.56 | 2.81× |
| 4v4 | 22.43 ± 3.30 | 88.36 ± 16.56 | 3.94× |
| 6v6 | 13.07 ± 1.18 | 51.21 ± 12.33 | 3.92× |
| 8v8 | 9.63 ± 1.24 | 42.01 ± 12.06 | 4.36× |
| 10v10 | 8.46 ± 1.61 | 55.43 ± 38.71 | 6.55× |
- JSBSim-based flight-dynamics modeling [9]
- Air-to-air missile, loadout, launch, and engagement-outcome simulation
- Configurable red and blue teams, initial states, weapons, rule-based opponents, and reward terms
- Gymnasium-style observation and action spaces
- Pure-Python backend for inspection, debugging, and rapid validation
- C++ backend for performance-intensive training and simulation
- Tacview-compatible
.acmireplay output - Built-in lightweight PPO training helper in
bvr_sim_rl - Web, OpenGL, and DX11 visualization paths
- skrl: the repository includes the lightweight
bvr_sim_rl/PPO entry point for installation checks, training, and evaluation. See the skrl project. - MARLBenchmark / off-policy:
rl_envs/env_marlbenchmark.pyprovides per-agent observations, centralized shared observations, and legacy Gym space interfaces. See the upstream marlbenchmark/off-policy framework and the lizi-Margin/off-policy branch maintained for this project. - HARL:
rl_envs/env_harl.pyimplements HARL's local/shared observations, rewards, done flags, information dictionaries, and available-action tuple. See PKU-MARL/HARL and the corresponding paper, Heterogeneous-Agent Reinforcement Learning [1]. - UHRL / HMP2G:
rl_envs/env_wrapper.pytargets the internal UHRL branch and implementsScenarioConfig,make_env, team mapping, state, and available-action interfaces. UHRL is a private branch developed from the public HMP2G framework; see Unreal-MAP [2] for the public background and application of HMP2G.
MARLBenchmark and UHRL support here means that framework interfaces are implemented; it does not imply validation of every algorithm provided by those frameworks. All adapters can select either the Python or C++ simulation backend.
Run all commands below from the repository root.
Python 3.8 or later is required; Python 3.10 or later is recommended.
python -m venv .venvWindows PowerShell:
.\.venv\Scripts\Activate.ps1Linux/macOS:
source .venv/bin/activatepip install -e ".[rl]"If the current shell does not accept the quoted extras syntax, install in two steps:
pip install -e .
pip install skrl torchThe Python backend requires no native-extension build. Use the tested PPO entry point and the Python-specific convergence configuration:
python -m bvr_sim_rl \
--backend python \
--config scripts/experiments/ppo/f16_1v1_convergence_python.jsonc \
--timesteps 100000 \
--num-envs 8 \
--rollouts 256 \
--seed 1111 \
--logdir runs/bvr_sim_rl/python_f16_1v1_seed1111Monitor the run in another terminal:
tensorboard --logdir runs/bvr_sim_rl/python_f16_1v1_seed1111timesteps counts vector steps. With 8 independent environments, this command collects approximately 800,000 environment transitions.
The C++ backend is intended for longer, higher-throughput training. Build the native extension first.
Windows:
bvr_sim\build_windows.batLinux:
bash bvr_sim/build_linux.shThen use the tested C++ training entry point and convergence configuration:
python -m bvr_sim_rl \
--backend cpp \
--config scripts/experiments/ppo/f16_1v1_convergence_cpp.jsonc \
--timesteps 80000 \
--num-envs 8 \
--rollouts 256 \
--seed 1111 \
--logdir runs/bvr_sim_rl/cpp_f16_1v1_seed1111Monitor reward, episode return, and episode length in another terminal:
tensorboard --logdir runs/bvr_sim_rl/cpp_f16_1v1_seed1111The 80,000-step command collects approximately 640,000 environment transitions and takes about one hour on the tested machine. In TensorBoard, Info / return_per_step is the most useful convergence signal when episode lengths are changing.
Pure-Python backend:
import json
from bvr_sim import BVR3DEnv
with open("scripts/tests/demo_config.json", "r", encoding="utf-8") as f:
config = json.load(f)
env = BVR3DEnv(config, logdir="./test_logs")
obs, info = env.reset()
done = False
while not done:
obs, reward, dones, info = env.step({})
done = info["episode_done"]C++ backend:
import commentjson
from bvr_sim import BVR3DEnvCpp
with open("scripts/tests/demo_config_cpp.jsonc", "r", encoding="utf-8") as f:
config = commentjson.load(f)
env = BVR3DEnvCpp(config, rl_index=[0], log_file_path="./test_logs/bvr_sim.log")
obs, info = env.reset()
done = False
while not done:
action = env.action_space.sample().reshape(1, -1)
obs, reward, dones, info = env.step(action)
done = info["episode_done"]The lightweight training module is located at:
bvr_sim_rl/
It creates either BVR3DEnv or BVR3DEnvCpp, converts simulator observations into tensors accepted by skrl, and quantizes the continuous actions produced by PPO back into the simulator's discrete actions.
The simulator's native action space is:
MultiDiscrete([15, 15, 9, 2])
The four dimensions represent:
delta_heading, delta_altitude, delta_speed, shoot
The PPO side uses a continuous action space:
Box(-1, 1, shape=(4,))
See docs/rl/ppo_quickstart.md for details.
Python backend smoke test:
python scripts/tests/test_py.pyC++ backend smoke test:
python scripts/tests/test_cpp.pyC++ unit tests:
python scripts/tests/cpp_unit_tests.pyRun the complete test suite, which rebuilds the native extension:
python scripts/run_tests.pySome tests and examples generate Tacview-compatible replays at:
test_logs/*.acmi
Open the .acmi files with Tacview to inspect an engagement.
The C++ backend provides several real-time visualization paths:
- Web visualization:
EmbeddedWebServer+bvr_sim/web/ - Native OpenGL viewer
- Native Windows DX11 game mode
After building the C++ backend, run:
python scripts/experimental/run_web_viz.py
python scripts/experimental/run_opengl_viz.py
python scripts/experimental/run_dx11_viz.pyDemo Video · Quick Start · PPO Training · Configuration · RL Integration
- Getting Started
- PPO Training Guide
- Installation
- Configuration
- External RL Framework Integration
- Developer Architecture
.github/
media/ README demo videos, thumbnails, and paper figures
bvr_sim/ Main Python package and C++ sources
src_py/ Python simulation, rewards, and rule-based opponents
src_cxx/ C++ simulation core
web/ Web visualization frontend
bvr_sim_rl/ Lightweight skrl PPO training entry point
scripts/ Development and evaluation entry points
benchmarks/ Benchmark configurations and scripts
experimental/ Experimental scenarios and visualization examples
experiments/paper/ Paper experiment scripts and configurations
tests/ Smoke tests and unit tests
docs/ Documentation, design notes, and usage guides
Do not edit generated directories manually:
bvr_sim/build/
bvr_sim/install/
test_logs/
benchmark_logs/
runs/
BVR Sim is released under GPLv3. See LICENSE.
For third-party dependencies and bundled-resource notices, see docs/legal/THIRD_PARTY_LICENSES.md.
- Yifan Zhong, Jakub Grudzien Kuba, Xidong Feng, Siyi Hu, Jiaming Ji, and Yaodong Yang. “Heterogeneous-Agent Reinforcement Learning.” Journal of Machine Learning Research, 25(32):1–67, 2024. Paper
- Tianyi Hu, Qingxu Fu, Zhiqiang Pu, Yuan Wang, and Tenghai Qiu. “Unreal-MAP: Unreal-Engine-Based General Platform for Multi-Agent Reinforcement Learning.” Proceedings of the AAAI Conference on Artificial Intelligence, 40(35):29486–29494, 2026. Paper · DOI
- Edvards Scukins, Markus Klein, Lars Kroon, and Petter Ögren. “BVR Gym: A Reinforcement Learning Environment for Beyond-Visual-Range Air Combat.” arXiv preprint arXiv:2403.17533, 2024. Paper · DOI
- Qihan Liu, Yuhua Jiang, and Xiaoteng Ma. “Light Aircraft Game: A Lightweight, Scalable, Gym-Wrapped Aircraft Competitive Environment with Baseline Reinforcement Learning Algorithms.” GitHub repository, 2022. Repository
- Hanzhong Cao. “Light Aircraft Game: Basic Implementation and Training Results Analysis.” arXiv preprint arXiv:2506.14164, 2025. Paper
- André R. Kuroswiski, Annie S. Wu, and Angelo Passaro. “B-ACE: An Open Lightweight Beyond Visual Range Air Combat Simulation Environment for Multi-Agent Reinforcement Learning.” 2024 Interservice/Industry Training, Simulation and Education Conference (I/ITSEC), Paper No. 24464, 2024.
- Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. “The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games.” Advances in Neural Information Processing Systems, 35:24611–24624, 2022. DOI
- Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang. “Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning.” International Conference on Learning Representations, 2022. Paper
- Jon S. Berndt. “JSBSim: An Open Source Flight Dynamics Model in C++.” AIAA Modeling and Simulation Technologies Conference and Exhibit, AIAA Paper 2004-4923, 2004.


