Skip to the page
GUIDES/TRAIN AN RL POLICY

Train an RL policy

By the end of this page you will have a Gymnasium environment over a simulated market, an episode played under a random policy and under fixed ones, and an episode you can replay exactly from its seed. The spaces, the reward and every argument are in the reference, Gym environment.

The code blocks on this page are one Python session and run in order. They use no training library. env is a standard Gymnasium environment, so a library that trains on those can take it.

Install the rl extra

tradefloor.gym needs numpy and gymnasium, which the rl extra installs.

pip install "tradefloor[rl]"

Build the environment

TradingEnv takes a universe and a seed and builds a market for each episode. An episode here is 5 trading days of 6 steps each.

import numpy as np
import tradefloor as tf
from tradefloor.gym import TradingEnv

universe = tf.Universe.random(5, seed=101)
env = TradingEnv(universe=universe, seed=42, days=5)
obs, info = env.reset()
print(env.observation_space)
print(env.action_space)
print(info)
Box(-inf, inf, (11,), float64)
Box(-1.0, 1.0, (5,), float64)
{'seed': 42, 'model_fingerprint': 'pt-v20'}

The observation holds each name's log return since the last step, each holding as a fraction of net worth, and cash as a fraction of net worth. The action is a target weight per name. The environment passes Gymnasium's env_checker. It warns that the observation space is unbounded, and that it cannot test render modes on an environment not built by gymnasium.make.

Play an episode

A random policy samples the action space. Seed the space, or the policy draws different actions on every run.

env.action_space.seed(0)
total = 0.0
done = False
while not done:
    obs, reward, terminated, truncated, info = env.step(env.action_space.sample())
    total += reward
    done = terminated or truncated
print(f"{info['step']} steps, total reward {total:+,.0f}")
print(f"net worth {info['net_worth']:,.0f}, leverage {info['leverage']:.2f}, "
      f"refused {info['rejected']}, scaled {info['scaled']}")
30 steps, total reward +6,543
net worth 1,006,543, leverage 1.95, refused 0, scaled True

The reward is the step's change in net worth, in dollars, so the rewards of an episode add up to its P&L. info after a step holds the book: net worth, cash, leverage, the trades refused, and whether the action was scaled down to fit the leverage cap. An episode is terminated when net worth reaches zero and truncated after its last step.

Your trades move prices

The policy trades in the market it observes. Its orders take liquidity from the book, and the prices it sees next include that. Here two fixed policies play the same seed with ten million in cash. One holds nothing, the other puts all of net worth into AAA.

def play(env, policy, seed):
    # one episode: AAA's log return at each step, and the total reward
    obs, info = env.reset(seed=seed)
    aaa, total, done = [], 0.0, False
    while not done:
        obs, reward, terminated, truncated, info = env.step(policy(obs))
        aaa.append(obs[0])
        total += reward
        done = terminated or truncated
    return np.array(aaa), total

flat = lambda obs: np.zeros(5)                    # hold nothing
all_in = lambda obs: np.array([1.0, 0, 0, 0, 0])  # everything in AAA

big = TradingEnv(universe=universe, seed=42, days=5, cash=10_000_000)
quiet, _ = play(big, flat, seed=42)
loud, _ = play(big, all_in, seed=42)
print(f"AAA's first step: {quiet[0] * 1e4:+.1f} bps holding nothing, "
      f"{loud[0] * 1e4:+.1f} bps buying")
print(f"AAA over the episode: {np.expm1(quiet.sum()):+.2%} and "
      f"{np.expm1(loud.sum()):+.2%}")
AAA's first step: +2.1 bps holding nothing, +9.3 bps buying
AAA over the episode: +0.93% and +0.98%

The buy moved AAA's first step from +2.1 to +9.3 bps, and AAA ends the episode higher than in the market where nobody bought it. With the default million in cash the same policy moves that first step by under 1 bp. The reward is measured after the market has moved, so it already includes what the policy's own trading cost, and a policy that trades large learns in a market that responds to it.

Replay an episode

The package version, the preset, the universe and the seed fix the market. The same actions on the same seed give the same rewards, and reset(seed=info["seed"]) replays any episode.

first, reward_a = play(env, all_in, seed=7)
again, reward_b = play(env, all_in, seed=7)
print(reward_a == reward_b, np.array_equal(first, again))

obs, info = env.reset()
print("next seed", info["seed"])
True True
next seed 4058335883

A reset() without a seed draws the next seed from the environment's own generator, so a training loop meets a new market each episode and the same sequence of markets on every run. Record the seeds a policy trained on and someone else can rebuild those episodes.

Next steps