Skip to content
lizi-MarginPublic

About

BVR air combat game for RL

Resources

Contributing

Security policy

Stars

20 stars

Watchers

0 watching

Forks

Repository files navigation

BVR Sim: A Fast RL Environment for Heterogeneous Air-Combat

English | 简体中文

Python 3.8+ C++ backend Gymnasium compatible PPO with skrl License GPLv3

BVR-Sim is a 3D beyond-visual-range air-combat environment for reinforcement-learning research, with Python and accelerated C++ backends and a built-in skrl PPO training pipeline for immediate experimentation.

Heterogeneous F-22/F-16 6v6 Human–AI Real-Time Air Combat Heterogeneous Air Combat + Ground Air Defense
Heterogeneous F-22 and F-16 6-vs-6 engagement Human and AI real-time air-combat game Heterogeneous air combat and ground air-defense reinforcement learning
Well-Tuned IPPO: 2v2 Composite Missile Guidance + FDM UnrealCV + UE5 Real-Time Rendering
Well-tuned IPPO in a 2-vs-2 engagement Composite missile guidance with flight dynamics UnrealCV and Unreal Engine 5 real-time rendering

BVR Sim game mode cockpit view

System Architecture

BVR Sim system architecture

The shared tactical interface maps heading, altitude, speed, and fire commands to aircraft-specific inner-loop controllers. Simulation, learning interfaces, and optional telemetry or rendering remain decoupled, allowing the same scenario to run through the Python reference backend or accelerated C++ backend.

Supported Aircraft

The C++ JSBSim backend currently provides mappings for the following aircraft families:

  • F-15
  • F-16
  • F/A-18
  • F-22
  • F-4N Phantom II
  • AJ 37 and JA 37 Viggen

Different aircraft models, controller parameters, sensors, and weapon loadouts can coexist in the same scenario.

Environment Comparison

The following comparison reproduces Table 1 of the paper for the cited public releases. A cross means that documented or direct support was not found in the corresponding release.

Capability BVR Gym [3] LAG [4, 5] B-ACE [6] BVR Sim
Open source ✓ ✓ ✓ ✓
JSBSim aircraft dynamics ✓ ✓ × ✓
BVR-class missile engagement ✓ × ✓ ✓
Mixed aircraft models in one scenario × × × ✓
Built-in maneuver-and-fire interface Maneuver only ✓ ✓ ✓
Interchangeable Python/native C++ backends × × × ✓
Uniform fixed-width entity table × × × ✓
Rule baseline and self-play ✓ ✓ ✓ ✓
ACMI/Tacview export or telemetry ✓ ✓ × ✓
Real-time 3D visualization FlightGear Tacview Godot Native viewers

MARL Validation

HAPPO and MAPPO mean episode reward on the BVR Sim 2v2 task

The archived HAPPO [8] and MAPPO [7] traces above reproduce Fig. 3 of the paper and verify end-to-end integration on the same 2-vs-2 BVR task. Each curve is from one run, so the figure is an integration check rather than an algorithm ranking.

Performance

BVR Sim provides two complementary flight-dynamics paths for different stages of reinforcement-learning development.

JSBSim FDM: Well-Tuned PPO 1v1 Simple FDM: Rapid RL Validation
Well-tuned PPO with JSBSim flight dynamics in a 1-vs-1 engagement Rapid reinforcement-learning validation with Simple FDM

The JSBSim FDM path provides higher-fidelity, near-realistic aircraft dynamics for flight-control and engagement studies. Simple FDM is a lightweight alternative for fast RL pipeline validation, reward debugging, and early convergence checks.

Table 3 of the paper reports headless throughput at a 0.4-second decision interval. Values are mean ± standard deviation over three repeats.

Scale Python steps/s C++ steps/s Speedup
1v1 95.84 ± 8.07 260.65 ± 2.75 2.72×
2v2 51.01 ± 8.09 143.58 ± 4.56 2.81×
4v4 22.43 ± 3.30 88.36 ± 16.56 3.94×
6v6 13.07 ± 1.18 51.21 ± 12.33 3.92×
8v8 9.63 ± 1.24 42.01 ± 12.06 4.36×
10v10 8.46 ± 1.61 55.43 ± 38.71 6.55×

Core Features

  • JSBSim-based flight-dynamics modeling [9]
  • Air-to-air missile, loadout, launch, and engagement-outcome simulation
  • Configurable red and blue teams, initial states, weapons, rule-based opponents, and reward terms
  • Gymnasium-style observation and action spaces
  • Pure-Python backend for inspection, debugging, and rapid validation
  • C++ backend for performance-intensive training and simulation
  • Tacview-compatible .acmi replay output
  • Built-in lightweight PPO training helper in bvr_sim_rl
  • Web, OpenGL, and DX11 visualization paths

Supported Reinforcement-Learning Frameworks

  • skrl: the repository includes the lightweight bvr_sim_rl/ PPO entry point for installation checks, training, and evaluation. See the skrl project.
  • MARLBenchmark / off-policy: rl_envs/env_marlbenchmark.py provides per-agent observations, centralized shared observations, and legacy Gym space interfaces. See the upstream marlbenchmark/off-policy framework and the lizi-Margin/off-policy branch maintained for this project.
  • HARL: rl_envs/env_harl.py implements HARL's local/shared observations, rewards, done flags, information dictionaries, and available-action tuple. See PKU-MARL/HARL and the corresponding paper, Heterogeneous-Agent Reinforcement Learning [1].
  • UHRL / HMP2G: rl_envs/env_wrapper.py targets the internal UHRL branch and implements ScenarioConfig, make_env, team mapping, state, and available-action interfaces. UHRL is a private branch developed from the public HMP2G framework; see Unreal-MAP [2] for the public background and application of HMP2G.

MARLBenchmark and UHRL support here means that framework interfaces are implemented; it does not imply validation of every algorithm provided by those frameworks. All adapters can select either the Python or C++ simulation backend.

Quick Start: Train PPO Directly

Run all commands below from the repository root.

1. Create a Virtual Environment

Python 3.8 or later is required; Python 3.10 or later is recommended.

python -m venv .venv

Windows PowerShell:

.\.venv\Scripts\Activate.ps1

Linux/macOS:

source .venv/bin/activate

2. Install the Project and Training Dependencies

pip install -e ".[rl]"

If the current shell does not accept the quoted extras syntax, install in two steps:

pip install -e .
pip install skrl torch

3. Train with the Python Backend

The Python backend requires no native-extension build. Use the tested PPO entry point and the Python-specific convergence configuration:

python -m bvr_sim_rl \
  --backend python \
  --config scripts/experiments/ppo/f16_1v1_convergence_python.jsonc \
  --timesteps 100000 \
  --num-envs 8 \
  --rollouts 256 \
  --seed 1111 \
  --logdir runs/bvr_sim_rl/python_f16_1v1_seed1111

Monitor the run in another terminal:

tensorboard --logdir runs/bvr_sim_rl/python_f16_1v1_seed1111

timesteps counts vector steps. With 8 independent environments, this command collects approximately 800,000 environment transitions.

4. Train with the C++ Backend

The C++ backend is intended for longer, higher-throughput training. Build the native extension first.

Windows:

bvr_sim\build_windows.bat

Linux:

bash bvr_sim/build_linux.sh

Then use the tested C++ training entry point and convergence configuration:

python -m bvr_sim_rl \
  --backend cpp \
  --config scripts/experiments/ppo/f16_1v1_convergence_cpp.jsonc \
  --timesteps 80000 \
  --num-envs 8 \
  --rollouts 256 \
  --seed 1111 \
  --logdir runs/bvr_sim_rl/cpp_f16_1v1_seed1111

Monitor reward, episode return, and episode length in another terminal:

tensorboard --logdir runs/bvr_sim_rl/cpp_f16_1v1_seed1111

The 80,000-step command collects approximately 640,000 environment transitions and takes about one hour on the tested machine. In TensorBoard, Info / return_per_step is the most useful convergence signal when episode lengths are changing.

Minimal Python Example

Pure-Python backend:

import json
from bvr_sim import BVR3DEnv

with open("scripts/tests/demo_config.json", "r", encoding="utf-8") as f:
    config = json.load(f)

env = BVR3DEnv(config, logdir="./test_logs")
obs, info = env.reset()

done = False
while not done:
    obs, reward, dones, info = env.step({})
    done = info["episode_done"]

C++ backend:

import commentjson
from bvr_sim import BVR3DEnvCpp

with open("scripts/tests/demo_config_cpp.jsonc", "r", encoding="utf-8") as f:
    config = commentjson.load(f)

env = BVR3DEnvCpp(config, rl_index=[0], log_file_path="./test_logs/bvr_sim.log")
obs, info = env.reset()

done = False
while not done:
    action = env.action_space.sample().reshape(1, -1)
    obs, reward, dones, info = env.step(action)
    done = info["episode_done"]

PPO Training Entry Point

The lightweight training module is located at:

bvr_sim_rl/

It creates either BVR3DEnv or BVR3DEnvCpp, converts simulator observations into tensors accepted by skrl, and quantizes the continuous actions produced by PPO back into the simulator's discrete actions.

The simulator's native action space is:

MultiDiscrete([15, 15, 9, 2])

The four dimensions represent:

delta_heading, delta_altitude, delta_speed, shoot

The PPO side uses a continuous action space:

Box(-1, 1, shape=(4,))

See docs/rl/ppo_quickstart.md for details.

Validate the Simulator

Python backend smoke test:

python scripts/tests/test_py.py

C++ backend smoke test:

python scripts/tests/test_cpp.py

C++ unit tests:

python scripts/tests/cpp_unit_tests.py

Run the complete test suite, which rebuilds the native extension:

python scripts/run_tests.py

Replay Output

Some tests and examples generate Tacview-compatible replays at:

test_logs/*.acmi

Open the .acmi files with Tacview to inspect an engagement.

Visualization

The C++ backend provides several real-time visualization paths:

  • Web visualization: EmbeddedWebServer + bvr_sim/web/
  • Native OpenGL viewer
  • Native Windows DX11 game mode

After building the C++ backend, run:

python scripts/experimental/run_web_viz.py
python scripts/experimental/run_opengl_viz.py
python scripts/experimental/run_dx11_viz.py

Documentation Navigation

Demo Video · Quick Start · PPO Training · Configuration · RL Integration

Documentation Index

Project Structure

.github/
  media/                     README demo videos, thumbnails, and paper figures
bvr_sim/                     Main Python package and C++ sources
  src_py/                    Python simulation, rewards, and rule-based opponents
  src_cxx/                   C++ simulation core
  web/                       Web visualization frontend
bvr_sim_rl/                  Lightweight skrl PPO training entry point
scripts/                     Development and evaluation entry points
  benchmarks/               Benchmark configurations and scripts
  experimental/             Experimental scenarios and visualization examples
  experiments/paper/        Paper experiment scripts and configurations
  tests/                    Smoke tests and unit tests
docs/                        Documentation, design notes, and usage guides

Do not edit generated directories manually:

bvr_sim/build/
bvr_sim/install/
test_logs/
benchmark_logs/
runs/

License

BVR Sim is released under GPLv3. See LICENSE.

For third-party dependencies and bundled-resource notices, see docs/legal/THIRD_PARTY_LICENSES.md.

References

  1. Yifan Zhong, Jakub Grudzien Kuba, Xidong Feng, Siyi Hu, Jiaming Ji, and Yaodong Yang. “Heterogeneous-Agent Reinforcement Learning.” Journal of Machine Learning Research, 25(32):1–67, 2024. Paper
  2. Tianyi Hu, Qingxu Fu, Zhiqiang Pu, Yuan Wang, and Tenghai Qiu. “Unreal-MAP: Unreal-Engine-Based General Platform for Multi-Agent Reinforcement Learning.” Proceedings of the AAAI Conference on Artificial Intelligence, 40(35):29486–29494, 2026. Paper · DOI
  3. Edvards Scukins, Markus Klein, Lars Kroon, and Petter Ögren. “BVR Gym: A Reinforcement Learning Environment for Beyond-Visual-Range Air Combat.” arXiv preprint arXiv:2403.17533, 2024. Paper · DOI
  4. Qihan Liu, Yuhua Jiang, and Xiaoteng Ma. “Light Aircraft Game: A Lightweight, Scalable, Gym-Wrapped Aircraft Competitive Environment with Baseline Reinforcement Learning Algorithms.” GitHub repository, 2022. Repository
  5. Hanzhong Cao. “Light Aircraft Game: Basic Implementation and Training Results Analysis.” arXiv preprint arXiv:2506.14164, 2025. Paper
  6. André R. Kuroswiski, Annie S. Wu, and Angelo Passaro. “B-ACE: An Open Lightweight Beyond Visual Range Air Combat Simulation Environment for Multi-Agent Reinforcement Learning.” 2024 Interservice/Industry Training, Simulation and Education Conference (I/ITSEC), Paper No. 24464, 2024.
  7. Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. “The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games.” Advances in Neural Information Processing Systems, 35:24611–24624, 2022. DOI
  8. Jakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen, Fanglei Sun, Jun Wang, and Yaodong Yang. “Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning.” International Conference on Learning Representations, 2022. Paper
  9. Jon S. Berndt. “JSBSim: An Open Source Flight Dynamics Model in C++.” AIAA Modeling and Simulation Technologies Conference and Exhibit, AIAA Paper 2004-4923, 2004.

About

BVR air combat game for RL

Resources

Contributing

Security policy

Stars

20 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages