Train an RL policy
By the end of this page you will have a Gymnasium environment over a simulated market, an episode played under a random policy and under fixed ones, and an episode you can replay exactly from its seed. The spaces, the reward and every argument are in the reference, Gym environment.
The code blocks on this page are one Python session and run in order. They use no training library. env is a standard Gymnasium environment, so a library that trains on those can take it.
Install the rl extra
tradefloor.gym needs numpy and gymnasium, which the rl extra installs.
pip install "tradefloor[rl]"
Build the environment
TradingEnv takes a universe and a seed and builds a market for each episode. An episode here is 5 trading days of 6 steps each.
import numpy as np import tradefloor as tf from tradefloor.gym import TradingEnv universe = tf.Universe.random(5, seed=101) env = TradingEnv(universe=universe, seed=42, days=5) obs, info = env.reset() print(env.observation_space) print(env.action_space) print(info)
Box(-inf, inf, (11,), float64)
Box(-1.0, 1.0, (5,), float64)
{'seed': 42, 'model_fingerprint': 'pt-v20'}The observation holds each name's log return since the last step, each holding as a fraction of net worth, and cash as a fraction of net worth. The action is a target weight per name. The environment passes Gymnasium's env_checker. It warns that the observation space is unbounded, and that it cannot test render modes on an environment not built by gymnasium.make.
Play an episode
A random policy samples the action space. Seed the space, or the policy draws different actions on every run.
env.action_space.seed(0)
total = 0.0
done = False
while not done:
obs, reward, terminated, truncated, info = env.step(env.action_space.sample())
total += reward
done = terminated or truncated
print(f"{info['step']} steps, total reward {total:+,.0f}")
print(f"net worth {info['net_worth']:,.0f}, leverage {info['leverage']:.2f}, "
f"refused {info['rejected']}, scaled {info['scaled']}")30 steps, total reward +6,543 net worth 1,006,543, leverage 1.95, refused 0, scaled True
The reward is the step's change in net worth, in dollars, so the rewards of an episode add up to its P&L. info after a step holds the book: net worth, cash, leverage, the trades refused, and whether the action was scaled down to fit the leverage cap. An episode is terminated when net worth reaches zero and truncated after its last step.
Your trades move prices
The policy trades in the market it observes. Its orders take liquidity from the book, and the prices it sees next include that. Here two fixed policies play the same seed with ten million in cash. One holds nothing, the other puts all of net worth into AAA.
def play(env, policy, seed):
# one episode: AAA's log return at each step, and the total reward
obs, info = env.reset(seed=seed)
aaa, total, done = [], 0.0, False
while not done:
obs, reward, terminated, truncated, info = env.step(policy(obs))
aaa.append(obs[0])
total += reward
done = terminated or truncated
return np.array(aaa), total
flat = lambda obs: np.zeros(5) # hold nothing
all_in = lambda obs: np.array([1.0, 0, 0, 0, 0]) # everything in AAA
big = TradingEnv(universe=universe, seed=42, days=5, cash=10_000_000)
quiet, _ = play(big, flat, seed=42)
loud, _ = play(big, all_in, seed=42)
print(f"AAA's first step: {quiet[0] * 1e4:+.1f} bps holding nothing, "
f"{loud[0] * 1e4:+.1f} bps buying")
print(f"AAA over the episode: {np.expm1(quiet.sum()):+.2%} and "
f"{np.expm1(loud.sum()):+.2%}")AAA's first step: +2.1 bps holding nothing, +9.3 bps buying AAA over the episode: +0.93% and +0.98%
The buy moved AAA's first step from +2.1 to +9.3 bps, and AAA ends the episode higher than in the market where nobody bought it. With the default million in cash the same policy moves that first step by under 1 bp. The reward is measured after the market has moved, so it already includes what the policy's own trading cost, and a policy that trades large learns in a market that responds to it.
Replay an episode
The package version, the preset, the universe and the seed fix the market. The same actions on the same seed give the same rewards, and reset(seed=info["seed"]) replays any episode.
first, reward_a = play(env, all_in, seed=7)
again, reward_b = play(env, all_in, seed=7)
print(reward_a == reward_b, np.array_equal(first, again))
obs, info = env.reset()
print("next seed", info["seed"])True True next seed 4058335883
A reset() without a seed draws the next seed from the environment's own generator, so a training loop meets a new market each episode and the same sequence of markets on every run. Record the seeds a policy trained on and someone else can rebuild those episodes.
Next steps
- Spaces and reward defines the observation, the action, the leverage cap and every
infokey. - TradingEnv lists every argument, including
model=for another preset. - Limits of a trained policy says what a policy trained here can and cannot tell you about a real market.