Skip to the page
API REFERENCE/GYM ENVIRONMENT

Gym environment

tradefloor.gym.TradingEnv is a Gymnasium environment, the standard Python interface for reinforcement learning. The policy's orders trade in the market's book and move its prices, and each seed is a new market.

TradingEnv

TradingEnv(
    *,
    universe: Sequence[Instrument],
    seed: int = 0,
    macro: Macro | None = None,
    days: int = 5,
    steps_per_day: int = 6,
    ticks_per_step: int = 65,
    cash: float = 1_000_000.0,
    max_leverage: float | None = 2.0,
    start: tuple[int, int, int] = (9, 30, 3),
    model: str | ModelParams | None = None,
    trusted_agents: bool = False,
)
ArgumentDefaultMeaning
universerequiredThe roster. Its order is part of the market.
seed0The first episode's market, and the seed of the generator that draws later episodes' seeds.
macroNoneThe economy on day 0. None uses Macro()'s defaults.
days5Trading days in an episode.
steps_per_day6Steps a day. An episode is days * steps_per_day steps.
ticks_per_step65Simulated minutes the market runs after each action.
cash1000000.0Starting cash.
max_leverage2.0Cap on gross exposure as a multiple of net worth. None removes it.
start(9, 30, 3)Hour, minute and day of the week the first session opens on.
modelNoneA preset name or a ModelParams, the same for every episode. None is the default preset, pt-v20.
trusted_agentsFalseTrue makes env.engine and env.portfolio the live objects. By default they are read-only views.

It needs numpy and gymnasium, from pip install "tradefloor[rl]". It passes Gymnasium's env_checker, which warns only that the observation space is unbounded.

Spaces and reward

For a roster of n names:

SpaceHolds
ObservationBox(-inf, inf, (2n + 1,), float64)Each name's log return since the previous step, each holding as a fraction of net worth, then cash as a fraction of net worth.
ActionBox(-1, 1, (n,), float64)A target weight per name, as a fraction of net worth. Negative is short.

The observation space is unbounded because no finite bound holds: a step's return is limited only by the circuit breaker on each of its ticks, and cash goes negative when the book is levered. An action outside [-1, 1] is clipped. When the absolute weights add up to more than the leverage cap allows, every weight is scaled by one factor to a gross exposure of max_leverage / (1 + 0.01 * max_leverage), 1.96 under the default cap of 2, and info["scaled"] is True. The step sells what it shrinks before it buys what it grows. A trade the book or the cap refuses is counted in info["rejected"], and the episode goes on.

The reward is the step's change in net worth, in dollars: net worth after the step's session less net worth before it. It is measured after the market has moved, so it includes the cost of the policy's own trading. A step's fills reach the market once, on the first tick of the session that follows, so rewards for a policy that trades differ from 0.8.1 and earlier. An episode is terminated when net worth reaches zero and truncated after its last step.

info keyReturned byMeaning
seedresetThe seed the episode ran. reset(seed=info["seed"]) replays it.
model_fingerprintresetThe preset's name, or custom- and a hash for changed coefficients.
trustedresetTrue, present only under trusted_agents=True.
net_worth, cashstepNet worth and cash after the step.
leveragestepGross exposure as a multiple of net worth.
rejectedstepTrades refused this step.
scaledstepWhether the action was scaled down to fit the leverage cap.
stepstepSteps taken in the episode.

An episode

import numpy as np
import tradefloor as tf
from tradefloor.gym import TradingEnv

universe = tf.Universe.random(5, seed=101)
env = TradingEnv(universe=universe, seed=42, days=5)

obs, info = env.reset()
print(obs.shape, env.action_space.shape, info["seed"])

pnl = 0.0
done = False
while not done:
    action = np.full(5, 0.6)       # 60% of net worth in each name: 3x gross
    obs, reward, terminated, truncated, info = env.step(action)
    pnl += reward
    done = terminated or truncated
print(info["step"], info["scaled"], round(info["leverage"], 2), round(pnl))
(11,) (5,) 42
30 True 1.97 38961

The action asks for 3 times net worth, so every step scales it to fit the cap of 2.

Seeds and episodes

The first reset() runs the constructor's seed, and reset(seed=n) runs seed n. Each later reset() without a seed draws a new seed below 2**32 from the environment's generator, so a loop of 1,000 resets meets 1,000 markets, the same 1,000 on every run. To see how much of a score came from the market, hold the universe fixed and change the seed. The package version, the preset, the universe's fingerprint and the seed fix an episode, so someone else can replay the episodes a policy trained on.

Limits of a trained policy

A policy that learns the model

Prices come from a known model, and a policy is very good at finding that model's structure. A high score can mean the policy found the herding term, a property of the model that no real market has. Compare the trained policy with buy-and-hold across many seeds, and before you make a claim about a real market, read How it is measured, which shows where the model matches real markets.

Weaker volatility memory than real markets

A volatile period fades faster in the model than in a real market. On pt-v20, volatility clustering, the tendency of volatile days to follow volatile days, is below real at every lag: about a quarter of the real strength one day apart and a sixth twenty days apart. A policy that sizes its trades on a volatility estimate of one month or longer learns less memory than real markets have.