# tradefloor > A reproducible evaluation environment for financial AI agents: a > deterministic market simulator with order-book execution, macro dynamics > and causal ground truth. > Version 0.8.7. Every page, in reading order. ====================================================================== # Quickstart https://docs.tradefloor.dev/ Install tradefloor, run a five-company market for 20 days, score a simple trading agent on it and read its scorecard, in one script you can download. ====================================================================== GETTING STARTED/ QUICKSTART Quickstart tradefloor is a simulated stock market with a limit order book and an economy, for testing trading strategies and AI agents. This page installs it, runs a small market, scores a simple agent on that market and explains the result, in about five minutes. The snippets on this page run in order in one Python session, and each one builds on the last. A block labeled Output shows what the code above it prints on tradefloor 0.8.7 with the default preset. To run the whole page at once, download quickstart.py and run python quickstart.py. Install pip install tradefloor tradefloor needs Python 3.11 or later and installs no other packages. Install lists the platforms with prebuilt wheels and the extras that some features need. import tradefloor as tf print(tf.__version__, tf.preset_record()["preset"]) 0.8.7 pt-v20 The first value is the package version. The second is the default preset, the named and frozen set of model coefficients a market runs on when you do not choose one. Run a small market import struct universe = tf.Universe.random(5, seed=11) # five made-up companies engine = tf.Engine(seed=42, universe=universe) # a market on the default preset engine.run_days(20) # run it for 20 trading days bars = engine.bars(grain="day") # one row per company per day print(bars.num_rows, "daily bars:", ", ".join(bars.columns)) closes = struct.unpack("<5d", engine.prices()) # each company's price now for company, close in zip(universe, closes): print(f"{company.ticker} {company.sector:<22} " f"{company.initial_price:7.2f} -> {close:7.2f}") 100 daily bars: day, bar, instrument_id, open, high, low, close, volume AAA technology 135.37 -> 133.98 AAB financial_services 282.62 -> 290.24 AAC healthcare 28.59 -> 29.34 AAD energy 12.73 -> 13.42 AAE consumer_discretionary 31.22 -> 32.41 The universe is the list of companies the market trades, in order. Universe.random(5, seed=11) makes five companies with made-up tickers, sectors and fundamentals, and it makes the same five every time. The engine's seed, 42, fixes every random draw the market takes, so this block prints the same numbers on every run. The first line of output counts 100 daily bars: one for each of the 5 companies on each of the 20 days, each with an open, high, low, close and volume. Each line after it is one company: its ticker, its sector, its price on day zero and its price when the 20 days ended. AAA, a technology company, started at 135.37 and ended at 133.98. bars is an Arrow table, which pandas, polars, pyarrow and duckdb read without copying. engine.prices() returns each company's last price as little-endian float64 bytes in roster order, and struct.unpack("<5d", ...) reads the five of them with no extra package. engine.truth() gives each company's fair value and the factors behind every price move. Engine and data lists every table. Score a simple agent An agent is any object with an act(obs) method. The harness calls it at every decision step, six a day by default, and act returns the orders to send: a dict of ticker to a number of shares, positive to buy and negative to sell, or None to send nothing. class EqualWeight: """Put about 100,000 into each company at the first step, then hold.""" def act(self, obs): if obs.step == 0: return {ticker: int(100_000 / price) # a number of shares for ticker, price in zip(obs.tickers, obs.prices)} return None # hold scores = tf.evaluate({"equal": EqualWeight()}, seed=42, universe=universe, days=20) print(scores["equal"]) Scorecard('equal', pnl=13,036, return=+1.30%, trades=5, impact=+0.38bps, sharpe=+1.62, vol=10.4%, in_market=100%, exposure=0.51x) tf.evaluate builds its own market from the seed and universe you pass, gives the agent 1,000,000 in cash and runs it for 20 days of six steps each. The orders fill against the simulated order book, so every trade pays the spread and moves the price a little. The quantities are shares, and a value such as 0.2 buys a fifth of one share (shares and portfolio weights). Understand the result Field Here What it means pnl 13,036 Net worth at the end minus the 1,000,000 the agent started with. return +1.30% The same profit as a percentage of the starting cash. trades 5 Fills. Each of the five market orders here, one per company, filled at once and counts once. A limit order counts once for each part that fills. impact +0.38bps How far the agent's own trades moved the prices of what it traded, against the same market with nobody trading, weighted by the money traded. Positive means the move cost the agent. sharpe +1.62 The mean daily return over its standard deviation, annualized over 252 days, with no risk-free rate subtracted. Under 20 days it prints as n/a (short run). vol 10.4% The annualized volatility of the daily returns. in_market 100% The share of steps that ended with a position open. exposure 0.51x The average of all positions, long and short, as a multiple of net worth. 1.0 is fully invested. The scorecard adds errors=N when a step raised or the market refused an order, and the lines are in scores["equal"].errors. Rejected orders explains each kind of line and what to change. A profit on its own says little, so score a reference agent on the same market: scores = tf.evaluate({"equal": EqualWeight(), "buy_and_hold": tf.baselines.BuyAndHold()}, seed=42, universe=universe, days=20) print(scores["buy_and_hold"]) print({name: round(gap) for name, gap in tf.versus_buy_and_hold(scores).items()}) Scorecard('buy_and_hold', pnl=23,765, return=+2.38%, trades=5, impact=+0.43bps, sharpe=+1.63, vol=19.3%, in_market=100%, exposure=0.94x) {'equal': -10729} Each agent trades its own copy of the same starting market. Buy-and-hold put nearly all its cash into the five companies (exposure 0.94x) and made 23,765. The equal-weight agent put in about half (0.51x) and made 13,036, so versus_buy_and_hold puts it 10,729 behind. This is one market on one seed, and another seed can reverse the order, so compare agents across seeds before you draw a conclusion. Reproducing a run Run this page again on tradefloor 0.8.7 and every number comes out the same. The market depends on the universe, the economy on day zero, the seed, the preset and the package version. It also depends on the agent's orders, because orders move prices. Reproducibility lists each condition, says which ones hold across platforms and releases, and shows how to save a run as a manifest that someone else can replay. Next steps Compare agents across seeds One seed is one sample. tf.rank runs the same agents on many seeds and counts, seed by seed, which one did better, as Compare strategies shows. Fork a market Copy a running market, change one input in one copy and run both on. The copies share their past exactly, so a later difference comes from the change. Fork a market walks through it. Connect an LLM Run an agent built with the OpenAI Agents SDK, PydanticAI, LangGraph or your own model call, and record its answers so the run replays without the model. LLM adapters shows how. Train an RL policy tradefloor.gym wraps a market as a Gymnasium environment, after pip install "tradefloor[rl]", and Gym environment documents it. If something on this page fails, Troubleshooting lists the usual causes, and the Glossary defines the terms used here. ====================================================================== # Install https://docs.tradefloor.dev/install.html Install tradefloor with pip, check the version and default preset, and add the optional extras for MCP, Gymnasium, Arrow tables and LLM agent frameworks. ====================================================================== GETTING STARTED/ INSTALL Install pip install tradefloor The package has no required dependencies. Everything that needs another library, such as the MCP server or the Gymnasium environment, is an optional extra. Requirements tradefloor needs CPython 3.11 or later. Each release is tested on 3.11, 3.12 and 3.13, and ships prebuilt wheels for five targets: Operating system Architectures Linux x86_64, aarch64 macOS arm64 (Apple silicon), x86_64 (Intel) Windows x86_64 On any other platform pip builds the package from source, which needs a Rust toolchain. The same engine is a Rust crate, added with cargo add tradefloor. Every release runs the same fixed simulations on all five targets before it is published and stops if any result differs, so a seed gives the same market on each of them. Reproducibility says what that check covers and what it leaves out. Checking the install import tradefloor as tf print(tf.__version__) print(tf.preset_record()["preset"]) 0.8.7 pt-v20 The first line is the package version. The second is the default preset, the set of model coefficients a market runs on when you do not name one. Optional extras Name an extra in square brackets. Quote the argument, because some shells read the brackets themselves: pip install "tradefloor[mcp]" Extra Installs Needed for arrow pyarrow reading a run's output tables with pyarrow rl NumPy and Gymnasium tradefloor.gym, the gym environment mcp the MCP SDK and pyarrow the tradefloor-mcp server (Local MCP server) openai-agents the OpenAI Agents SDK the OpenAI Agents adapter, in live mode pydantic-ai pydantic-ai-slim the PydanticAI adapter, in live mode langgraph LangGraph the LangGraph adapter, in live mode finrobot FinRobot and its dependencies the FinRobot adapter, in live mode. Python 3.11 only claude the Anthropic SDK and Pydantic the example examples/08-claude-agent.py The output tables are Arrow tables, so polars, pandas and duckdb read them too, and you need the arrow extra only if you want pyarrow itself. The framework adapters need their framework only to call a model. Replaying a recorded run needs none of them, so a plain pip install tradefloor replays a recording. In live mode, an adapter whose framework is missing raises MissingDependencyError, an ImportError whose message gives the pip install command for the extra. Pinning a version The API can change in any release before 1.0, so pin the version you tested with: pip install "tradefloor==0.8.7" A pinned version also pins the default preset. Support policy says what a patch release may change and how long a release line gets fixes. Examples and notebooks The numbered examples and notebooks are in the GitHub repository, and the package does not include them. Clone the repository to run them. Some of them read recorded runs from its tests/fixtures/ folder. If an install or an import fails, see Troubleshooting. ====================================================================== # Reproducibility https://docs.tradefloor.dev/reproducibility.html What fixes a tradefloor market, what an untraded and a traded run promise across platforms and releases, and what a preset name does and does not pin. ====================================================================== GETTING STARTED/ REPRODUCIBILITY Reproducibility A tradefloor market is a function of its inputs. Run it twice with the same inputs on the same release and you get the same market to the last bit, on any of the five platforms tradefloor ships wheels for. The inputs These fix a market. Change any one of them and you get a different market. The universe: the companies, in order. The engine draws random numbers in roster order, so the same companies in another order are a different market. universe.fingerprint, a sha256 of the roster, covers the order as well as the companies. The economy on day zero: the macro_state passed to tf.Engine. Leaving it out is an input too. The engine then settles its own default opening, which is a different market from passing tf.Macro() with the same values. The seed: the seed passed to tf.Engine, any integer from 0 to 2**64 - 1. Universe.random takes a separate seed, which picks the companies and not the market. The preset: the model passed to tf.Engine or tf.evaluate. Leaving it out uses the installed version's default, which is pt-v20 in 0.8.7. The package version, which decides what the default preset is and which code runs. Any scenario applied to the market, day by day. For a traded run, the agent's orders, in the order they were sent, and for an LLM agent the answers the model gave. The snippets on this page run in order in one Python session. This one runs the same 20-day market with one input changed at a time and compares a digest of the whole engine state: import tradefloor as tf universe = tf.Universe.random(5, seed=11) def final_state(roster=universe, seed=42, macro=None, model=None): """Run 20 days and return a digest of the whole engine state.""" engine = tf.Engine(seed=seed, universe=roster, macro_state=macro, model=model) engine.run_days(20) return engine.state_hash() reference = final_state() print("same inputs again: ", final_state() == reference) print("roster reversed: ", final_state(roster=tf.Universe(reversed(universe))) == reference) print("VIX 30 on day zero: ", final_state(macro=tf.Macro(vix=30.0)) == reference) print("seed 43: ", final_state(seed=43) == reference) print("preset pt-v19: ", final_state(model="pt-v19") == reference) same inputs again: True roster reversed: False VIX 30 on day zero: False seed 43: False preset pt-v19: False Markets with and without an agent A market with no agent orders in it depends on the inputs above and nothing else. The same inputs give the same market, bit for bit, on Linux, macOS and Windows. tradefloor carries its own exp, log, pow, sin and cos, so the platform's math library cannot change a result, and every release runs fixed simulations on all five wheel targets and stops if any digest differs. A run with an agent in it has one more input. The agent's orders fill against the order book and move prices, and the moved prices are what the agent sees at its next decision. So the agent's decisions are part of the input, and two agents on the same seed trade two different markets after their first trade. A Python agent that decides the same way from the same observations gives the same run again on the same release. Anything in it that varies, such as an unseeded random number or a value read from the clock, changes its orders and so the market. Its own arithmetic is outside tradefloor's checks too. Python 3.12 changed how sum() adds floats, so an agent that sums prices with sum() can send different orders on 3.11 and 3.12. An LLM agent can answer differently every time it is asked. The LLM adapters record every call and its answer in a transcript, and a replay feeds those answers back without calling the model. A replay looks each answer up by a digest of the exact observation the model was shown, so it stops with ReplayMiss when anything upstream has changed. Replay mismatches lists the messages. Every input an engine consumes, orders included, is in engine.order_log, and tf.replay(log, seed=..., universe=...) rebuilds the market from that log without the agent's code. Across releases A shipped preset never changes. Its coefficients are fixed when it first ships in a tagged release, and a better coefficient becomes a new preset with a new name. So an untraded market on a named preset replays exactly in every later release, and every shipped preset from pt-v1 on can still be selected with model="pt-v12" or similar. Each release from 0.8.5 on checks one known-answer digest per shipped preset on all five targets. The default preset moves between releases. pt-v20 became the default in 0.8.5, replacing pt-v19, so a run that left model out on 0.8.1 gives a different market on 0.8.5 and later. Name the preset in anything you publish. Traded runs have a narrower promise. In 0.8.5 and later an agent's fills reach the market once, on the next tick, where 0.8.1 fed them in as order flow on every tick of the next step. That change applies to every preset, so a traded run recorded before 0.8.5 matches up to its first trade and differs after it. Each release also checks one traded run on all five targets: the reference agents through tf.evaluate on pt-v20, with their orders, fills and scorecards. Runs with your own agents go through the same code. Within the 0.8 long-term support line a patch release has to leave every one of these digests unchanged (Support policy). Preset names and preset values A preset name is fixed only in tagged releases. A development build or an edited source tree can carry other values under the same name. A RunManifest and a transcript both store the values and refuse a mismatch, and a coefficient changed through the API, as in tf.ModelParams.from_preset("pt-v20", momentum_theta=0.05), fingerprints as custom- and eight hex digits, never as the preset. To pin a model, install a tagged release from PyPI or crates.io and name the preset. Saving a run A RunManifest holds what a stranger needs to rebuild a market: the version, the preset and its coefficients, the seed, the universe, the economy on day zero, any scenario and the order log, with a digest of the market the run ended on. engine = tf.Engine(seed=42, universe=universe) engine.run_days(20) text = tf.RunManifest.of(engine, seed=42, universe=universe).to_json() rebuilt = tf.RunManifest.from_json(text).reproduce() # replays, then checks print(rebuilt.state_hash() == engine.state_hash()) True reproduce() checks every component against its fingerprint before it replays, and raises an error naming the part that disagrees. It first runs a small fixed probe simulation on the default preset, so a manifest written on a release with a different default (0.8.1, read on 0.8.5 or later) is refused. Rebuild such a run from its named preset, seed and universe instead. A manifest checks the market and carries no score. tf.evaluate and tf.rank write no manifest, so to let a reader check a score, publish the agent's code, the call that scored it and the seeds. Citing tradefloor lists what to report beside a result. ====================================================================== # Troubleshooting https://docs.tradefloor.dev/troubleshooting.html Fixes for install errors, missing extras, agents that never trade, refused orders, shares sent as weights, missing history, replay errors and model failures. ====================================================================== GETTING STARTED/ TROUBLESHOOTING Troubleshooting Error text is quoted as tradefloor 0.8.7 prints it, and numbers in it vary from run to run. Installation and extras No matching distribution Symptom pip install tradefloor ends with No matching distribution found for tradefloor. Cause The Python running pip is older than 3.11, the oldest version tradefloor supports. Fix Check with python --version, then install into a 3.11 or later environment, for example python3.12 -m venv .venv. A build from source Symptom pip downloads a .tar.gz instead of a wheel, and the build fails with an error about Rust, cargo or maturin. Cause There is no prebuilt wheel for your platform. Wheels cover Linux x86_64 and aarch64, macOS arm64 and x86_64, and Windows x86_64, and anything else builds from source. Fix Install a Rust toolchain from rustup.rs and run pip again, or use one of the five platforms (Install). Brackets in zsh Symptom pip install tradefloor[mcp] stops with zsh: no matches found: tradefloor[mcp]. Cause zsh reads the square brackets as a file pattern before pip sees them. Fix Quote the argument: pip install "tradefloor[mcp]". Missing extras Symptom An import or a call raises an ImportError that ends in a pip install command, such as tradefloor.gym needs numpy and gymnasium, and numpy is not installed. Install them with: pip install 'tradefloor[rl]'. Cause The feature needs an optional package that the core install leaves out. The MCP server, the gym environment and the live mode of each LLM adapter work this way. Fix Run the command in the message. Optional extras lists every extra. FinRobot on Python 3.12 and later Symptom pip install "tradefloor[finrobot]" stops with Could not find a version that satisfies the requirement finrobot>=0.1.5. Cause FinRobot supports Python 3.10 and 3.11, and tradefloor needs 3.11 or later, so the extra installs on 3.11 only. Fix Use a Python 3.11 environment for live FinRobot runs. Replaying a recorded FinRobot run needs no extra, on any supported Python. An agent that makes no trades The scorecard reads trades=0 and a P&L of 0. Look at the end of the scorecard first: errors=N there means the agent tried and something failed. Acting on the wrong step Symptom No errors, no trades, or trades only on the first day. Cause obs.step counts decision steps over the whole run, from 0 to six times the number of days, less one. A check written as obs.step == 0 to mean "every morning" is true once. obs.step_of_day starts again at 0 each day. Fix Use obs.is_first_step_of_day for once a day and obs.step == 0 for once a run. Errors on every step Symptom A warning such as Agent 'mine' failed on all 120 of its steps, so its score is empty. First error: ..., and errors=120 on the scorecard. Cause act raised, or returned something other than a mapping. A list of pairs gives act() must return a mapping of ticker to order, such as {'AAA': 100} or {'AAA': tf.Limit(100, 25.0)}, or None to trade nothing. It returned a list: .... Fix Print scores["mine"].errors, where each line names the step and what went wrong. A step that raised traded nothing, and the steps after it ran as normal. A class passed as the agent Symptom tf.evaluate raises before any market runs: Agent 'mine' is the class Mine, not an agent. Pass Mine() instead. Cause The agents mapping holds the class, and the harness needs an object with an act(obs) method. Fix Pass an instance: {"mine": Mine()}. Limit orders that never fill Symptom No errors, and trades stays at 0 while the agent sends tf.Limit orders. Cause A limit order waits in the book until the market reaches its price, and trades counts fills, so an order that never fills adds nothing. A buy limit well below the current price can wait for the whole run. A new limit on the same ticker replaces the one waiting there. Fix Set the limit price from obs.price(ticker), or send a plain number of shares for a market order that fills at once. An agent that needs past prices before it acts can also sit out a whole run, which Missing warm-up history covers. Rejected orders A refused order trades nothing, adds one to the scorecard's rejected and adds a line to errors, and the other orders from the same step still trade. Message in errors Cause Fix step 4: no instrument with ticker "aaa" in this universe The ticker is not in the market. Tickers are case-sensitive. Take tickers from obs.tickers. step 2: trade would take leverage to 2.37x, above the 2.00x limit The order would take gross positions past max_leverage times net worth, 2.0 by default. Counted again in leverage_refusals. Send a smaller order, or pass max_leverage to tf.evaluate. None removes the limit. step 0: the order for 'AAA' must be a number of shares, a tf.Limit or a tf.Cancel, got '100' (str) The quantity is a string or a bool. Send an int or a float. step 0: the order for 'AAB' must be finite, got nan A calculation produced NaN or infinity, often a division by zero. Check the inputs to the size calculation, and send nothing when they are missing. Once an agent's net worth reaches zero or below, the leverage limit refuses every order, because leverage over a net worth of zero is infinite. A market order larger than the book can absorb is not refused. It fills what the book holds, and the scorecard lists it in partial_fills, as in step 2: asked to buy 1,000,000,000,000 AAD; the book held 44,021,201, and the rest did not fill. Shares and portfolio weights Symptom The agent trades but the P&L is a few units of currency, with a warning such as Agent 'mine' asked for fractions of a share (0.2 of AAA, 0.2 of AAB, 0.2 of AAC, ...). act() returns numbers of shares, not portfolio weights. Cause act returned portfolio weights. {"AAA": 0.2} buys a fifth of one share of AAA. Fix Convert each weight to shares at the current price, as below. Each value act returns is a trade to make now, in shares, and not a position to hold. Returning {"AAA": 100} at every step buys 100 more shares at every step. To hold a target weight, trade the difference between the target and the current position: def act(self, obs): worth = obs.portfolio.net_worth() orders = {} for ticker, price in zip(obs.tickers, obs.prices): target = int(0.2 * worth / price) # 20% of net worth orders[ticker] = target - obs.position(ticker) return orders tf.baselines.rebalance(obs, {"AAA": 0.2, "AAB": 0.2}) does the same, skips trades of less than one share and caps each trade at 2% of the company's average daily volume, which limits what the agent pays in impact. Missing warm-up history Symptom An agent that needs, say, 20 days of prices trades nothing for the first 20 days, or nothing at all in a short run. obs.history.bars("AAA", last=20) returns fewer than 20 bars, and none at the first step. Cause tf.evaluate and tf.rank start at day 0 with no history unless asked for one. Fix Pass history_days=20. The market then runs 20 untraded days before day 0, and obs.history holds their daily bars at the first decision, labeled day -20 to day -1. A warm-up changes which days are scored. The scored days continue the warmed market, so on the same seed they are different days from a run without one, and the scorecard records history_days to say so. A warm-up can be up to 2520 days, ten 252-day years. The day in progress is never in obs.history, and each scored day joins it after its close. The reference agents and tf.StrategySpec strategies keep their own price history and ignore obs.history. While the history is still empty, obs.history.bars() returns an empty list even for a ticker the market does not list, so check tickers against obs.tickers. Replay mismatches Manifest refusals RunManifest.reproduce() raises a ValidationError naming the part that disagreed. Message starts with Cause Fix this build does not reproduce the manifest's era The manifest was written on a release whose default preset or arithmetic differs from this one. The message names the version and platform that wrote it. Install that version, or rebuild the run from its named preset, seed and universe. this manifest ran model preset 'pt-vN', which this build does not ship The preset is newer than this release, or came from a development build. Install a release that ships the preset. model preset 'pt-vN' disagrees between this manifest and this build Same name, different values: one side was a development build (Preset names and preset values). Reproduce on the build that wrote the manifest. the replay ran but did not rebuild the recorded market The inputs matched and the market did not. If the message says the draw counts differ, the order log does not cover everything that happened to the engine. Replay part of the log with tf.replay(log, ..., until=n) and find the first step that differs. the universe in this manifest does not match its recorded fingerprint The file was edited after it was written. Get an unedited copy. A traded run recorded before 0.8.5 replays up to its first trade and differs after it on 0.8.5 and later, because 0.8.5 changed how fills reach the market. Replay it on the release that recorded it. Refused LLM replays A replay looks up each answer by a digest of the exact observation the model was shown. When it cannot find one it raises ReplayMiss, which stops tf.evaluate with a note naming the step, the agent and the seed, so it is never scored as an agent that held cash. Message starts with Cause Fix no recorded response for step N (day D, digest ...) Something that feeds the observation changed after recording: the seed, the universe, history_days, the instructions or the market settings. A recording made before 0.8.5 can fail this way too, because 0.8.5 changed the observation the model is shown. Replay with exactly the call that recorded it, on the release that recorded it, or record again live. this transcript was recorded against a different simulation preset The replay runs another preset, often the default. Pass the recorded preset, as in tf.evaluate(..., model="pt-v19"). this transcript was recorded under observation payload version The recording was made by a release with another observation payload. Since 0.8.5 each recording names its payload version. Replay on the release named in the message. the recorded entry for step N (day D, digest ...) holds a null response The live call failed during recording, and the failure was recorded. Record the run again live. Model-provider failures Failed model calls Symptom The scorecard has errors=N and few or no trades, with lines such as step 12: FrameworkError: openai-agents raised APIConnectionError instead of returning a decision: .... Cause The call to the model never completed: no network, a missing or wrong API key, a rate limit or a timeout. The adapters add no retry of their own, and tf.evaluate scores the step as one that traded nothing and carries on. Fix Read the exception type at the end of the line, fix the key, quota or connection, and run again. Compare runs by their errors count as well as their P&L, because a failed call and a decision to hold both trade nothing. Unreadable decisions Symptom Lines in errors such as DecisionError: no JSON object in the framework response, no 'actions' key in the decision (keys present: ...) or the framework returned an empty response, so there is no decision to validate. Cause The model answered, and the answer was not a decision: an actions list and an optional rationale. The adapter refuses to guess what was meant. Fix Ask for that shape in the instructions, or use the framework's structured output. An answer of {"actions": []} is a valid decision to hold. Refused actions Symptom The decision traded, and errors and rejected also count one action, for example a symbol the market does not list. Cause Since decision schema 2 a bad action inside a good decision is refused on its own, and the rest of the decision trades. Fix Read the refused action: line in errors, which gives the reason. The PydanticAI request budget Symptom UsageLimitReached: the run at step N (day D) stopped on its request budget of ... Cause The PydanticAI adapter caps the requests one decision may make, at 8 by default. Fix Raise request_limit on the adapter, or pass None to remove the cap. Adapter warnings Message Cause Fix ... is starting a new run (day 0, step 0) but still holds N steps of prices from an earlier one One adapter ran a second evaluate or World, so its model sees the first market's prices in its returns. Build a new adapter for each run, as tf.rank does with its factory. this agent has output validators registered with @agent.output_validator PydanticAI refuses to change the output type of an agent with an output validator. Pass bind_output_type=False to keep your output type. tradefloor still validates the decision. GraphInterruptedError A LangGraph graph paused for a human. The market moves on as soon as act returns, so there is nothing to resume. Give the adapter a graph that decides without an interrupt. ====================================================================== # Glossary https://docs.tradefloor.dev/glossary.html Short definitions of the terms the tradefloor documentation uses, from preset, seed and universe to decision step, tick, fork, manifest and fair value. ====================================================================== GETTING STARTED/ GLOSSARY Glossary The terms these pages use in a specific sense, in alphabetical order. Each entry links to the page that covers it in full. Agent Any object with an act(obs) method. tf.evaluate, tf.rank and the other harnesses call it at every decision step, and it returns the orders to send as a dict of ticker to shares, or None to send nothing. An LLM agent is a model wrapped by one of the LLM adapters. Arrow table The format of a run's output tables, such as engine.bars() and engine.truth(). pandas, polars, pyarrow and duckdb read them without copying, and tradefloor depends on none of them. Value columns are float64, and a missing value is NaN. Basis point One hundredth of one percent, written bps. A trade's impact is reported in basis points. Decision step One call of an agent's act. tf.evaluate makes six a day by default (steps_per_day=6), each 65 ticks apart, so 20 days are 120 steps. obs.step counts steps over the whole run and obs.step_of_day starts again at 0 each day. Default preset The preset a market runs on when you do not name one. It is pt-v20 in tradefloor 0.8.7, since 0.8.5, and it can change between minor releases. Name the preset in a published result. Engine tf.Engine, the object that runs one market forward through time. It is built from a seed, a universe, the economy on day zero and a preset. Fair value What a company is worth on its fundamentals: its earnings or book value, priced at its sector's valuation and the interest rates of the day. The traded price is pulled back toward it. On pt-v20 news and market shocks also move fair value itself, mostly for good. engine.truth() reports it for every company at every tick, and How prices are made gives the model. Fingerprint A short hash that identifies one input. universe.fingerprint is a sha256 of the roster, order included. A model's fingerprint is the preset's name when its coefficients are exactly that preset's, and custom- with eight hex digits otherwise. A strategy built from a tf.StrategySpec has a fingerprint too, and a hand-written agent has none. Fork A copy of a running market that continues on its own. tf.branch(engine, 2) makes two copies in memory, which share their past bit for bit, so a difference between them afterwards comes from whatever you changed in one. A tf.Checkpoint saves a market to a file and rebuilds it later, in another process. Forks and counterfactuals covers both, with the calls for comparing two copies. Impact How far an agent's own trades moved the prices of what it traded, against the same market with nobody trading, in basis points. The scorecard weights it by the money traded in each company, and a positive value means the move cost the agent. Manifest A tf.RunManifest: one JSON file that records the package version, the preset and its coefficients, the seed, the universe, the economy on day zero, any scenario and the order log, with a digest of the finished market. reproduce() replays it and raises an error naming the part that disagrees, as Saving a run shows. Mispricing The log gap between a company's price and its fair value, the column mispricing_s in engine.truth(). The eleven factor columns beside it add up to its change at every tick. Order book The bids and offers waiting at each price for one company. An agent's orders fill against its depth, so a large order fills at worse prices than a small one and moves the price. Order log engine.order_log, the list of every input the engine has consumed, agent orders included. tf.replay(log, seed=..., universe=...) rebuilds the market from it. Preset A named, frozen set of model coefficients, such as pt-v20. Choose one with model="pt-v19" on tf.Engine or tf.evaluate. A shipped preset never changes, so a market on a named preset replays in later releases. Reproducibility says what a name pins, and Why pt-v20 lists every preset. Reference agents The five agents tf.reference_agents() returns, to score your own against on the same market: buy_and_hold, random, momentum, mean_reversion and oracle. tf.baselines.BuyAndHold buys every company in equal weight at its first decision step and never trades again, and tf.versus_buy_and_hold(scores) gives each agent's P&L minus buy-and-hold's. The oracle reads the model's hidden fair value, and its scorecard says so. Release checks The known-answer tests. Each release runs fixed simulations on all five wheel targets, hashes the results and compares the hashes with each other and with the committed ones, and any difference stops the release. Support policy lists the runs. Scenario A file of changes to apply to a market on given days, such as a rate rise or an oil shock, kept apart from the knock-on effects you assume follow from it. Seven ship with the package, tf.Scenario.load("liquidity_crisis") loads one, and Scenarios covers writing your own. Scorecard What tf.evaluate returns for each agent: its P&L, return, trades (the number of fills), impact, risk figures and any errors. Understand the result explains each field. Seed The integer that fixes every random draw a market takes, from 0 to 2**64 - 1. tf.Engine(seed=...) seeds the market and Universe.random(n, seed=...) seeds the companies, and the two are separate. Every seed below 2**32 gives the market it gave before 0.8.5. Tick One minute of simulated trading time. A trading day has 390 ticks, and engine.truth() has one row per company per tick. Transcript A recording of an LLM agent's model calls and answers, written by the LLM adapters. A replay feeds the answers back without calling the model, so the run repeats exactly. Universe The list of companies a market trades, in order. tf.Universe.random(n, seed=...) makes made-up companies, and tf.Universe.from_edgar(...) builds them from SEC filings. The order is part of the input: the same companies in another order are a different market. Validated scope The conditions the realism measurements cover: runs of up to one year (252 trading days) on a roster balanced across sectors, on the default preset. tf.envelope.check() refuses a question outside it and says why. How it is measured gives the measurements. Warm-up history Days the market runs before day 0 with nobody trading, so an agent that needs past prices has them at its first decision. Ask for them with history_days=N on tf.evaluate or tf.rank, and read them from obs.history, as Missing warm-up history describes. ====================================================================== # Write an agent https://docs.tradefloor.dev/guide-agent.html Write a trading agent for tradefloor: the act(obs) method, the orders it returns, sizing to a target weight, what the agent can see and warm-up history. ====================================================================== GUIDES/ WRITE AN AGENT Write an agent An agent is any Python object with an act(obs) method. This guide builds one, sizes its orders, reads what it is allowed to see and gives it price history before its first decision. Every argument, default and return value is in the reference, Agents and evaluation. Every code block on this page runs on its own. A first agent At each decision step, six a day by default, the harness builds an observation, passes it to act and reads back a dict of ticker to order. tf.evaluate runs the agent on a market built from a seed and a universe, here 40 generated companies, and returns a scorecard per agent. import tradefloor as tf market = tf.Universe.random(40, seed=111) class Quoter: # bid a cent under the last price at each open, cancel what is left at the close def act(self, obs): if obs.is_first_step_of_day: bid = round(obs.price("AAA") - 0.01, 2) return {"AAA": tf.Limit(500, bid)} if obs.is_last_step_of_day: return {"AAA": tf.Cancel()} return None card = tf.evaluate({"quoter": Quoter()}, seed=7, universe=market, days=10)["quoter"] print(card) Scorecard('quoter', pnl=20,597, return=+2.06%, trades=39, impact=+1.47bps, sharpe=n/a (short run), vol=18.4%, in_market=100%, exposure=0.56x) trades counts fills. Ten limit orders made 39 of them, because a waiting order fills in parts as the price comes to it. Each agent starts with 1,000,000 in cash and may hold positions worth up to twice its net worth. obs.step counts decision steps across the whole run, so day 1 starts at step 6. A guard written if obs.step == 0 therefore fires once a run, and an agent that means once a day reads obs.is_first_step_of_day or obs.step_of_day instead. Orders from act act returns a dict of ticker to order, or None to trade nothing that step. A plain number is a market order for that many shares, positive to buy and negative to sell or short. tf.Limit(quantity, price) waits in the book at its price or better, and tf.Cancel() withdraws every waiting order on the ticker. A new tf.Limit on a ticker replaces the one waiting there. The full table, with what happens to a partial fill, is in What act returns. A bad entry is refused on its own and the rest of the dict still trades. Each refusal adds a line to the scorecard's errors, so read errors before the P&L: import tradefloor as tf market = tf.Universe.random(40, seed=111) class Careless: # a quantity sent as text, and a ticker the market does not list def act(self, obs): if obs.step == 0: return {"AAA": 100, "AAB": "100", "ZZZZ": 50} return None card = tf.evaluate({"careless": Careless()}, seed=7, universe=market, days=1)["careless"] print(card.trades, card.rejected) for line in card.errors: print(line) 1 2 step 0: the order for 'AAB' must be a number of shares, a tf.Limit or a tf.Cancel, got '100' (str) step 0: no instrument with ticker "ZZZZ" in this universe If an agent's orders never fill or are refused, An agent that makes no trades and Rejected orders list the usual causes. Sizing to a target weight Values in the dict are shares, never weights. To hold 20% of net worth in AAA, the agent works out the holding it wants, 0.2 × net worth ÷ price rounded toward zero to whole shares (int() does this for both signs). It then sends that target minus the shares it holds, minus the shares in orders still waiting on the ticker, counting a waiting buy as positive and a waiting sell as negative. 0.2 * net_worth / price alone is the target holding. Sent as an order while the agent already holds the name, it buys the whole position again. A waiting limit order counts because it can still fill: a market order sent for the full gap while a bid waits can fill twice. A new tf.Limit on the same ticker replaces the waiting one, so when the agent resends a limit it sizes it against what is held alone. The library accepts a fraction of a share, so rounding is the agent's choice. A whole number matches what a broker takes, and the hosted app refuses fractions. evaluate warns when every order in a step is below one share, which is the usual sign of weights sent as shares. import tradefloor as tf def shares_to_send(obs, ticker, weight): """Target holding, current holding, waiting orders and the order to send.""" # whole shares, rounded toward zero: floor() for a long target target = int(weight * obs.portfolio.net_worth() / obs.price(ticker)) held = obs.position(ticker) waiting = sum(o["remaining"] if o["side"] == "buy" else -o["remaining"] for o in obs.portfolio.open_orders() if o["ticker"] == ticker) return target, held, waiting, target - held - waiting class TwentyPercent: # hold 20% of net worth in AAA: bid for it at the first open, # top up at the market when the gap passes 25 shares def act(self, obs): target, held, waiting, order = shares_to_send(obs, "AAA", 0.2) if obs.step in (0, 1, 6, 7): print(f"day {obs.day} step {obs.step_of_day}: target {target}, " f"held {held:.0f}, waiting {waiting:.0f}, order {order:.0f}") if obs.step == 0: return {"AAA": tf.Limit(order, round(obs.price("AAA") * 0.997, 2))} if obs.is_last_step_of_day: return {"AAA": tf.Cancel()} return {"AAA": order} if abs(order) > 25 else None market = tf.Universe.random(40, seed=111) card = tf.evaluate({"twenty": TwentyPercent()}, seed=7, universe=market, days=5)["twenty"] print(card) day 0 step 0: target 997, held 0, waiting 0, order 997 day 0 step 1: target 987, held 0, waiting 997, order -10 day 1 step 0: target 962, held 0, waiting 0, order 962 day 1 step 1: target 968, held 962, waiting 0, order 6 Scorecard('twenty', pnl=-7,057, return=-0.71%, trades=2, impact=+0.97bps, sharpe=n/a (short run), vol=8.1%, in_market=80%, exposure=0.16x) On day 0 the bid 0.3% under the price waits all day and never fills. At step 1 the target is 987 shares and the bid still covers 997 of them, so the order is −10 and inside the band, and the agent sends nothing. Without the waiting term it would have bought 987 shares at the market on top of the bid. The day's last step cancels the bid, so on day 1 nothing waits and the agent buys its 962 shares at the market. One step later it holds them, and the gap is 6 shares. Shares and portfolio weights covers the usual sizing mistakes. The agent's view The observation holds what a trader inside the market could know: the roster and last prices, each name's order book and average daily volume, the agent's own positions, cash and waiting orders, daily bars and the published economic figures in obs.history, and a view of the market in obs.engine. What obs.engine is depends on how the agent is run: Access How it is granted obs.engine and obs.portfolio Scorecard says Ordinary The default Read-only views. The market view serves prices, the public columns, each book, the published macro figures and the yield curve, and which names have news today. It has no fair value, no factor attribution, no future macro path and no fork. nothing Privileged privileged = True on the agent The same read-only views, plus obs.hidden: every column, the factor attribution, the model's coefficients, the current macro state and economy, and the fundamentals. Still read-only, with no fork and no future path. uses_hidden_state Trusted trusted_agents=True on the call The live engine and portfolio, as every run handed them out before 0.8.5. An agent can fork the engine and run the copy ahead, read the macro path a scenario will set, and write to the market. trusted Asking the ordinary view for anything it does not serve raises tf.SandboxError, which evaluate records in errors. On pt-v20 the business cycle and GDP growth arrive late, as the agencies publish them. The Oracle baseline is privileged. Under any access the harness compares the engine before and after each call to act and scores an agent that changed or copied it tampered. That check catches accidents and is no security boundary: agent code runs in the harness's own process, so run code you did not write in a separate process. What the agent sees in the reference lists each view's members. Warm-up history Every run starts at day 0 with no price history, so a rule that needs 20 days of bars would sit in cash for 20 days. history_days=N runs the market for N days before day 0 with nobody trading, and obs.history holds those days at the first decision. import tradefloor as tf market = tf.Universe.random(40, seed=111) class Breakout: # once a day, buy 300 shares of a name trading above its 20-day high def act(self, obs): if not obs.is_first_step_of_day: return None orders = {} for ticker in obs.tickers: bars = obs.history.bars(ticker, last=20) if len(bars) == 20 and obs.position(ticker) == 0 \ and obs.price(ticker) > max(bar["high"] for bar in bars): orders[ticker] = 300 return orders card = tf.evaluate({"breakout": Breakout()}, seed=7, universe=market, days=30, history_days=20)["breakout"] print(card) Scorecard('breakout', pnl=23,386, return=+2.34%, trades=15, impact=+0.06bps, sharpe=+3.93, vol=5.0%, in_market=97%, exposure=0.43x, history_days=20) The warm-up changes the market: the scored days continue the warmed one, so on the same seed they are different days from a run without it, and the scorecard records history_days. It takes at most 2,520 days, ten years of 252. Missing warm-up history covers an agent that sits in cash because it has none. Next steps Compare strategies puts the agent beside buy-and-hold and ranks it across many seeds. LLM adapters runs an agent built in an LLM framework through the same act method. Agents and evaluation has the signatures, the scorecard's fields and how evaluate treats errors. ====================================================================== # Compare strategies https://docs.tradefloor.dev/guide-compare.html Compare trading agents in tradefloor: against buy-and-hold on one market, across many seeds with rank and a paired sign test, and by what their trades cost. ====================================================================== GUIDES/ COMPARE STRATEGIES Compare strategies One market is one sample. This guide puts an agent beside buy-and-hold on a single market, then on many markets with tf.rank, reads the paired result, and prices what its own trading cost. The signatures and every field of the results are in the reference, Agents and evaluation. Every code block on this page runs on its own. The agent they compare is the Quoter from Write an agent, repeated in each block, on a universe of 40 generated companies. One market and the baselines tf.reference_agents() builds five baselines: buy-and-hold, a random trader, momentum, mean-reversion and the Oracle, which reads the model's fair value. tf.evaluate gives each agent its own copy of the same market, so no agent uses up liquidity another expected, and tf.versus_buy_and_hold gives each agent's P&L minus buy-and-hold's. import tradefloor as tf market = tf.Universe.random(40, seed=111) class Quoter: # bid a cent under the last price at each open, cancel what is left at the close def act(self, obs): if obs.is_first_step_of_day: bid = round(obs.price("AAA") - 0.01, 2) return {"AAA": tf.Limit(500, bid)} if obs.is_last_step_of_day: return {"AAA": tf.Cancel()} return None agents = tf.reference_agents(seed=3) agents["quoter"] = Quoter() scores = tf.evaluate(agents, seed=7, universe=market, days=10) for name, gap in tf.versus_buy_and_hold(scores).items(): print(f"{name:15} {gap:+12,.0f}") random -41,359 momentum -36,407 mean_reversion -44,998 oracle -7,242 quoter +10,121 On pt-v20 compare with buy-and-hold. A shock there moves fair value for good, so the Oracle's P&L follows the market's month and sets no ceiling, and tf.capture_ratio returns nothing and says why. To put several agents in one market, where one agent's trades move the prices another meets, use a World. Across seeds with rank tf.rank runs the same agents on many seeds, each a different market on the same universe. It calls the factory once per seed, so every seed gets new agents and no state carries from one market to the next. import tradefloor as tf market = tf.Universe.random(40, seed=111) class Quoter: # bid a cent under the last price at each open, cancel what is left at the close def act(self, obs): if obs.is_first_step_of_day: bid = round(obs.price("AAA") - 0.01, 2) return {"AAA": tf.Limit(500, bid)} if obs.is_last_step_of_day: return {"AAA": tf.Cancel()} return None def entrants(): # new agents for every seed, so no state carries from one market to the next agents = tf.reference_agents(seed=3) agents["quoter"] = Quoter() return agents ranking = tf.rank(entrants, seeds=range(8), universe=market, days=10) print(ranking.report()) print(ranking.separation("quoter", "buy_and_hold")) 8 seeds on universe 9be68b9bc37e... under model pt-v20 quoter vs buy_and_hold +7,064 a seed ahead 5/8 wins 5/8 buy_and_hold the benchmark median pnl +1,709 wins 2/8 mean_reversion vs buy_and_hold -23,726 a seed ahead 1/8 wins 1/8 random vs buy_and_hold -30,372 a seed ahead 0/8 wins 0/8 momentum vs buy_and_hold -48,869 a seed ahead 0/8 wins 0/8 The Oracle 'oracle' ran and has no row, because it is not a ceiling on this model (below). Its P&L over buy_and_hold's was -595 a seed, ahead on 5 of 8 seeds. No capture ratio on pt-v20. Market moves there mostly stick: each shock moves fair value for good, so even perfect knowledge of the model's fair value leaves little edge. The Oracle made money in 10 of 14 test markets and its P&L follows the market's month, so a fraction of it would measure the month, not the agent. Compare against buy-and-hold instead. {'a': 'quoter', 'b': 'buy_and_hold', 'wins_a': 5, 'wins_b': 3, 'ties': 0, 'paired_seeds': 8, 'decisive': False, 'p_value': 0.7265625} Reading the result Each row of the report is one agent against the benchmark, buy-and-hold by default, measured seed by seed on the same market. The figure before "a seed" is the mean of the agent's P&L minus buy-and-hold's, in currency, and on pt-v20 it is the figure to quote. "ahead" counts the seeds on which the agent earned more than buy-and-hold, and "wins" the seeds on which it had the highest P&L of every ranked agent. separation(a, b) is a paired sign test between two agents. The Quoter was ahead on 5 of 8 seeds and behind on 3. At 5 wins in 8 and p = 0.73, this run can't tell them apart, because a fair coin lands 5 to 3 or further from even that often. decisive is true only when one agent won on every seed, and it is a count, never a significance test. Report the paired count beside any pooled P&L, and add seeds before calling a winner. ranking.records["quoter"] holds the per-seed figures behind the row, described in AgentRecord. A row can also carry REFUSED, UNUSABLE or RAISED lines, which count the steps an agent's orders were refused, its act returned something other than a dict, or its code raised. Read those before the P&L. Ranking lists every member of the result. Execution cost tf.tca.analyse runs the same seed twice, with and without the agent's orders, and prices every fill against the market where it never traded. Real data cannot supply that untraded market, which is why arrival price and VWAP stand in for it there. import tradefloor as tf market = tf.Universe.random(40, seed=111) class RoundTrip: # buy 2,000 AAA, sell them back three steps later def act(self, obs): if obs.step == 0: return {"AAA": 2000} if obs.step == 3: return {"AAA": -2000} return None ex = tf.tca.analyse(RoundTrip(), seed=42, universe=market, days=1) print(round(ex.shortfall_bps(), 2)) 8.41 A positive shortfall is a cost. The engine has no slippage formula: a large order walks the book and pays for the depth it takes, and the cost of size follows the square-root law real metaorders follow (How it is measured, row C9). tca.analyse takes market orders only and refuses a tf.Limit, because the part of a limit order that waits fills inside a session, where the untraded run has no price to compare it with. tca.analyse and Execution have the arguments and the result. Next steps Fork a market changes one input in a copy of a market and compares the two futures. Record and replay makes a comparison that involves a model reproducible. ====================================================================== # Fork a market https://docs.tradefloor.dev/guide-fork.html Fork a tradefloor market to run two futures from one past: checkpoints, engine branches, and a counterfactual that runs one agent under a scenario and without it. ====================================================================== GUIDES/ FORK A MARKET Fork a market A fork copies a running market so you can run two futures from one past. Everything before the fork is identical in both copies, bit for bit, so a difference that appears afterwards comes from the input you changed. This guide forks a bare market, saves the point it forked from, and then runs a counterfactual with an agent in it. The signatures are in the reference, Forks and counterfactuals. Every code block on this page runs on its own. Two futures from one past tf.Checkpoint.of marks a point in a running engine, with its seed and universe, and branch builds independent engines at that point. Change one input in one copy and run both on. import struct import tradefloor as tf def last_price(engine, ticker): # prices() is each last price as little-endian float64 bytes, in roster order prices = struct.unpack(f"<{len(engine.tickers)}d", engine.prices()) return prices[engine.index_of(ticker)] universe = tf.Universe.random(40, seed=7) engine = tf.Engine(seed=42, universe=universe) engine.run_days(60) # one shared past mark = tf.Checkpoint.of(engine, universe=universe, seed=42) calm, hiked = mark.branch(2) # two engines at day 60 print("identical at the fork:", calm.state_hash() == hiked.state_hash()) hiked.pin_macro(corporate_bond_yield=0.09) # one change, in one copy calm.run_days(20) hiked.run_days(20) print(f"AAA after 20 more days: {last_price(calm, 'AAA'):.2f} calm, " f"{last_price(hiked, 'AAA'):.2f} with credit at 9%") text = mark.to_json() # outlives the process again = tf.Checkpoint.from_json(text).resume() print("resumed from text:", again.state_hash() == engine.state_hash()) identical at the fork: True AAA after 20 more days: 471.65 calm, 327.47 with credit at 9% resumed from text: True The state hash covers prices, every column, the random generators and the economy, so equal hashes mean the two copies are the same market. Pinning the corporate bond yield at 9% is the only difference between the two futures, and AAA ends about 31% lower under it. Branch or checkpoint There are two ways to copy a market, and they trade speed for durability: Cost Survives the process tf.branch(engine, 2) under 1 ms no Checkpoint.resume() or Checkpoint.branch() seconds, growing with the order log yes tf.branch copies the engine's state directly: every column and each random stream's position. A checkpoint stores the order log and rebuilds the engine by replaying it, so a checkpoint is what to save and cite in a published result. A checkpoint records the universe's fingerprint and refuses to load against a universe that changed, because two universes can share tickers and differ in fundamentals. It also refuses to load on a build whose arithmetic differs from the one that wrote it, and the error names both builds. That digest changed in 0.8.0, so a checkpoint written by 0.7.x loads only on 0.7.x. A counterfactual with an agent A World holds the market, one agent, its portfolio and the macro path. Forking a world copies all of them, so the same agent, holding the same position, meets two markets that differ in one input. Here the agent buys 300 shares of AAA on the first step and holds them, and one arm gets the packaged liquidity_crisis scenario. import tradefloor as tf class BuyOnce: # buy 300 AAA on the run's first step, then hold def act(self, obs): return {"AAA": 300} if obs.step == 0 else None universe = tf.Universe.random(24, seed=7) world = tf.World(seed=7, universe=universe, agent=BuyOnce()) world.run(50) # 50 shared days control, stress = world.fork("control", "stress") stress.apply(tf.Scenario.load("liquidity_crisis")) started = tf.agree(control, stress) print("identical at the fork:", started.identical) control.run(80) stress.run(80) result = tf.compare(control, stress, agreement=started) first = result.as_dict()["divergence"] print("prices first differ on day", first["prices"] // first["steps_per_day"]) print(result.render()) identical at the fork: True prices first differ on day 100 control stress ---------------------------------------------------------- --- agent behaviour --- final gross exposure 0.10x 0.08x steps it traded on 0 0 orders sent 0 0 unusable responses 0 0 turnover $0 $0 --- execution --- trades filled 0 0 partial fills 0 0 refused trades 0 0 cost against arrival $0 $0 the same, in bps +0.00 bps +0.00 bps --- portfolio --- cash $852,894 $852,894 final value $952,768 $925,182 P&L since the fork $-54,385 $-81,971 return since the fork -5.40% -8.14% max drawdown since 6.28% 10.19% P&L since inception $-47,232 $-74,818 return since inception -4.72% -7.48% Reading the comparison agree checks that the two arms match in engine state, order books, portfolio, agent state and random-stream positions, and compare with agreement= states that in the result. The scenario's shocks start on its own day 50, and apply counts that from the day it is applied, so the arms run identically for 50 more days and first differ on day 100. compare finds that first step separately for the macro path, the agent's decisions, its orders, the prices and the portfolios. This agent never trades after the fork, so both arms show no orders and the same cash, and the whole difference in value is the scenario's effect on the 300 shares it holds. An agent that reacts to prices would show its decisions diverging too. An external agent, such as a model behind an LLM adapter, runs in a world the same way, and world.manifest() records the run as a manifest, as in Record and replay. The result describes this simulated market under the one change you made. A claim about a real market needs evidence from a real market. Next steps Scenarios builds what World.apply takes. World lists every method of a world, including several agents in one market. ====================================================================== # Record and replay https://docs.tradefloor.dev/guide-replay.html Record an LLM agent's answers in tradefloor and replay the run without the model: transcripts, what a replay checks, saving a recording and a manifest of the market. ====================================================================== GUIDES/ RECORD AND REPLAY Record and replay A seed replays a tradefloor market exactly. A model behind an API answers differently every time and costs money on every call. Recording what the model answered, and replaying the run from that recording, makes a result from an LLM agent reproducible and lets a reader check it without an API key. The adapters' reference is LLM adapters and MCP, and setting an adapter up is covered in LLM adapters. Every code block on this page runs on its own except the two lines in Saving a recording, which continue the recording built in Recording and replay. None of them calls a model, because a plain function stands in for one, and none needs an optional extra. Recording and replay An adapter in mode="live" with a recorder calls the model and writes each answer to the recorder, a transcript. The same adapter in mode="replay" with that recording as its transcript answers from it and never calls the model. import json import tradefloor as tf from tradefloor.integrations.callable import callable_agent from tradefloor.integrations.common import AdapterInfo, Transcript, digest PROMPT = "You manage a portfolio. Answer with one JSON decision." def ask_model(payload): # your model call goes here; this stand-in answers in text first = payload["assets"][0] order = {"symbol": first["symbol"], "side": "BUY", "quantity": 400, "limit_price": first["best_bid"]} return json.dumps({"actions": [order], "rationale": "bid at the touch"}) def to_decision(raw, payload): decision = json.loads(raw) for action in decision["actions"]: action["quantity"] = min(action["quantity"], 100) # your risk cap return decision info = AdapterInfo(framework="callable", instructions_digest=digest(PROMPT)) market = tf.Universe.random(12, seed=4242) recording = Transcript() live = callable_agent(ask_model, postprocess=to_decision, info=info, mode="live", recorder=recording) first = tf.evaluate({"m": live}, seed=4242, universe=market, days=5)["m"] replay = callable_agent(postprocess=to_decision, info=info, mode="replay", transcript=recording) again = tf.evaluate({"m": replay}, seed=4242, universe=market, days=5)["m"] print("replay matches", first.pnl == again.pnl) print(len(recording), "recorded answers") print({key: recording.meta[key] for key in ("model_preset", "observation_schema_version", "decision_schema_version")}) replay matches True 5 recorded answers {'model_preset': 'pt-v20', 'observation_schema_version': '1', 'decision_schema_version': '2'} An adapter asks its model every 6 decision steps by default, once a simulated day, so five days are five decisions and five recorded answers. ask_model returns the model's raw text and to_decision turns it into a decision. The transcript records the raw text, so the parsing and the risk cap in to_decision run again on every replay. Code inside ask_model after the model call is recorded as its output and does not run on replay. Replay checks Each recorded answer is keyed by a digest of the exact input the model was sent, never by a step number. A replay that reaches an input with no recorded answer stops with ReplayMiss and names the step, so a changed experiment cannot quietly reuse the answers given to an old one. Change What happens on replay Seed, universe or cadence The input differs, so the first lookup misses and the run stops with ReplayMiss. Preset Refused with ReplayMiss naming both presets before any lookup. Recording and replay must run on the same preset, set with model= on evaluate or World. Observation or decision schema version Refused before the first lookup, naming both versions. The model's instructions Refused with ValidationError when the adapter is built, if the recording carries an instructions digest. The callable adapter carries one only when its AdapterInfo does, as here. Instructions get their own check because most adapters send them outside the input the replay is keyed on, so a changed prompt would otherwise match every key. The LangGraph adapter renders its instructions into the input, so there a changed prompt is a lookup miss. import json import tradefloor as tf from tradefloor.integrations.callable import callable_agent from tradefloor.integrations.common import (AdapterInfo, ReplayMiss, Transcript, digest) PROMPT = "You manage a portfolio. Answer with one JSON decision." def ask_model(payload): first = payload["assets"][0] order = {"symbol": first["symbol"], "side": "BUY", "quantity": 100} return json.dumps({"actions": [order], "rationale": "add"}) info = AdapterInfo(framework="callable", instructions_digest=digest(PROMPT)) market = tf.Universe.random(12, seed=4242) recording = Transcript() live = callable_agent(ask_model, info=info, mode="live", recorder=recording) tf.evaluate({"m": live}, seed=4242, universe=market, days=5) # the same recording, replayed on another seed replay = callable_agent(info=info, mode="replay", transcript=recording) try: tf.evaluate({"m": replay}, seed=4243, universe=market, days=5) except ReplayMiss: print("seed 4243: ReplayMiss") # the same recording, under an edited prompt edited = AdapterInfo(framework="callable", instructions_digest=digest(PROMPT + " Be bold.")) try: callable_agent(info=edited, mode="replay", transcript=recording) except tf.ValidationError: print("edited prompt: ValidationError") seed 4243: ReplayMiss edited prompt: ValidationError evaluate lets ReplayMiss end the run rather than scoring it as an error, because a broken recording scored as an error would read as an agent that held cash from that step on. Replay mismatches gives the fix for each failure. Saving a recording A transcript is plain JSON. recording.save(path) writes it and Transcript.load(path) reads it back, and to_json() and Transcript.from_json(text) do the same with a string: recording.save("run-4242.json") recording = Transcript.load("run-4242.json") A recording made before 0.8.5 has no schema versions, and its first lookup misses on 0.8.5 or later because the payload it was keyed on changed. Replay it on the release that made it. Recording the market The transcript reproduces the model. The market is reproduced from its seed and its order log, and a manifest, RunManifest, carries both, with the package version, the preset and a digest the replay has to match. evaluate and rank write no manifest, so run the agent in a World to get one. import json import tradefloor as tf from tradefloor.integrations.callable import callable_agent from tradefloor.integrations.common import AdapterInfo, Transcript, digest PROMPT = "You manage a portfolio. Answer with one JSON decision." def ask_model(payload): first = payload["assets"][0] order = {"symbol": first["symbol"], "side": "BUY", "quantity": 100} return json.dumps({"actions": [order], "rationale": "add"}) info = AdapterInfo(framework="callable", instructions_digest=digest(PROMPT)) market = tf.Universe.random(12, seed=4242) recording = Transcript() agent = callable_agent(ask_model, info=info, mode="live", recorder=recording) world = tf.World(seed=4242, universe=market, agent=agent) world.run(5) manifest = world.manifest() text = manifest.to_json() # publish this beside the transcript engine = tf.RunManifest.from_json(text).reproduce() print("market reproduced", engine.state_hash() == world.engine.state_hash()) print(manifest.result["days"], "days,", len(recording), "recorded answers") market reproduced True 5 days, 5 recorded answers reproduce() rebuilds the engine from the manifest and raises ValidationError naming the component that disagreed. It checks the market and carries no score. RunManifest lists what a manifest holds. Publishing a result To let a reader check a result from an LLM agent, publish: the transcript, which replays the model's answers the manifest, which replays the market adapter.provenance(), which names the framework, the model, the instructions digest and the settings that change what a decision can be the agent's code and the call that ran it, with its seeds To show a score was not tuned to its seeds, publish tf.commit(seeds, salt) before the run, then the seeds and the salt after it. tf.reveal checks the three agree. Bringing an LLM agent covers the fixed seven-market battery that tf.fingerprint.fingerprint runs to tell two versions of an agent apart. Next steps LLM adapters sets up the OpenAI Agents SDK, PydanticAI, LangGraph and FinRobot adapters, which record and replay the same way. Replay and recording in the reference covers each adapter's instruction check and what a record entry holds. ====================================================================== # LLM adapters https://docs.tradefloor.dev/llm-adapters.html Run an LLM agent inside a tradefloor market on your machine with the OpenAI Agents SDK, PydanticAI, LangGraph, FinRobot or a function: setup, a first run, fixes. ====================================================================== GUIDES/ LLM ADAPTERS LLM adapters An adapter runs an agent built in an LLM framework inside a simulated market on your machine. tradefloor sends the framework a JSON observation, the framework answers with a decision, and tradefloor checks the decision, places its orders in the simulation and scores the run. Each adapter is an ordinary agent with an act method, so it runs under tf.evaluate, tf.rank and World. Three ways to connect a model, and which of them place orders LLM adapters, on this page, run the model inside a simulated market in your own Python process, and its decisions become orders in that simulation. The local MCP server gives an MCP client simulation and evaluation tools. No tool places an order. The hosted app, in beta, lets a model or a bot place orders in saved markets on app.tradefloor.dev over MCP, an HTTP API or an Alpaca-shaped API. Setup Each framework is an optional extra. tradefloor needs Python 3.11 or later. Framework Install Module Minimum version A plain Python function pip install tradefloor tradefloor.integrations.callable none OpenAI Agents SDK pip install "tradefloor[openai-agents]" tradefloor.integrations.openai_agents openai-agents 0.22 PydanticAI pip install "tradefloor[pydantic-ai]" tradefloor.integrations.pydantic_ai pydantic-ai-slim 2.36 LangGraph pip install "tradefloor[langgraph]" tradefloor.integrations.langgraph langgraph 1.2 FinRobot pip install "tradefloor[finrobot]" tradefloor.integrations.finrobot Python 3.11 exactly A replay of a recorded run needs none of the extras. A first run with no API key The callable adapter wraps any function that takes the observation payload and returns a decision. A function that answers without a model is the fastest way to check the loop before any key or extra is involved. import json import tradefloor as tf from tradefloor.integrations.callable import callable_agent def ask_model(payload): # stands in for a model call: the payload is what a model is shown first = payload["assets"][0] order = {"symbol": first["symbol"], "side": "BUY", "quantity": 100} return json.dumps({"actions": [order], "rationale": "add 100 a day"}) adapter = callable_agent(ask_model) market = tf.Universe.random(12, seed=4242) card = tf.evaluate({"model": adapter}, seed=4242, universe=market, days=5)["model"] print(card) print(len(adapter.record), "decisions") Scorecard('model', pnl=1,027, return=+0.10%, trades=5, impact=+0.09bps, sharpe=n/a (short run), vol=0.4%, in_market=100%, exposure=0.02x) 5 decisions Replace the body of ask_model with a call to your model, and parse its text into a decision with postprocess=, as in Recording and replay. The record in adapter.record holds one entry per decision: the payload, the exact input, the raw response, the validated decision and the orders. A first run in each framework Each block below runs on its own with no API key, because it hands the adapter an offline model the framework ships for testing. They need the extra, and were tested with openai-agents 0.23.1, pydantic-ai-slim 2.53.0 and langgraph 1.2.12 on tradefloor 0.8.6. PydanticAI TestModel answers every request with the arguments it is given. Pass your own model through the adapter's model= argument the same way. import tradefloor as tf from pydantic_ai import Agent from pydantic_ai.models.test import TestModel from tradefloor.integrations.pydantic_ai import PydanticAIAdapter # a scripted model: no API key, the same answer at every decision offline = TestModel(custom_output_args={ "actions": [{"symbol": "AAA", "side": "BUY", "quantity": 200}], "rationale": "add to AAA"}) pm = Agent(instructions="You manage a small portfolio.") adapter = PydanticAIAdapter(pm, model=offline) market = tf.Universe.random(12, seed=4242) card = tf.evaluate({"pm": adapter}, seed=4242, universe=market, days=5)["pm"] print(card) print(len(adapter.record), "decisions") Scorecard('pm', pnl=2,055, return=+0.21%, trades=5, impact=+0.18bps, sharpe=n/a (short run), vol=0.7%, in_market=100%, exposure=0.03x) 5 decisions For a live run, build the Agent with your provider's model, as PydanticAI's documentation describes, and leave out model=. Your agent's tools, deps and instructions are passed through unchanged, and the adapter binds the decision schema as the output type for each run. OpenAI Agents SDK ScriptedModel from agents.testing answers each call with a function of the input. payload_of reads back the observation the adapter sent. import json import tradefloor as tf from agents import Agent from agents.testing import ModelStep, ScriptedModel, assistant_message from tradefloor.integrations.openai_agents import OpenAIAgentsAdapter, payload_of def answer(call): # reads the same payload a real model is sent first = payload_of(call)["assets"][0] order = {"symbol": first["symbol"], "side": "BUY", "quantity": 200} return [assistant_message(json.dumps({"actions": [order], "rationale": "add"}))] offline = ScriptedModel([ModelStep.respond(answer)] * 5) pm = Agent(name="Portfolio Manager", instructions="You manage a small portfolio.") adapter = OpenAIAgentsAdapter(pm, mode="live", model=offline) market = tf.Universe.random(12, seed=4242) card = tf.evaluate({"pm": adapter}, seed=4242, universe=market, days=5)["pm"] print(card) print(len(adapter.record), "decisions") Scorecard('pm', pnl=2,055, return=+0.21%, trades=5, impact=+0.18bps, sharpe=n/a (short run), vol=0.7%, in_market=100%, exposure=0.03x) 5 decisions OpenAIAgentsAdapter defaults to mode="replay", so a live run names mode="live", and openai_agent(agent) is the same adapter with live as its default. For a live run, give the Agent a model your account can use, set OPENAI_API_KEY and leave out model=. The adapter turns the SDK's tracing off for each run it starts unless you pass tracing=True, because tracing is on by default in the SDK and sends traces to OpenAI. LangGraph The adapter takes any compiled graph. By default it sends the payload under observation and as messages, and reads the decision from a decision key, an actions key or the last message. from typing import Any, TypedDict import tradefloor as tf from langgraph.graph import END, START, StateGraph from tradefloor.integrations.langgraph import LangGraphAdapter class State(TypedDict, total=False): observation: dict[str, Any] decision: dict[str, Any] def decide(state: State) -> State: # a rule in place of a model: buy 200 of the first name when flat first = state["observation"]["assets"][0] if first["position"] > 0: return {"decision": {"actions": [], "rationale": "holding"}} order = {"symbol": first["symbol"], "side": "BUY", "quantity": 200} return {"decision": {"actions": [order], "rationale": "open"}} builder = StateGraph(State) builder.add_node("decide", decide) builder.add_edge(START, "decide") builder.add_edge("decide", END) adapter = LangGraphAdapter(builder.compile()) market = tf.Universe.random(12, seed=4242) card = tf.evaluate({"graph": adapter}, seed=4242, universe=market, days=5)["graph"] print(card) print(len(adapter.record), "decisions") Scorecard('graph', pnl=544, return=+0.05%, trades=1, impact=+0.04bps, sharpe=n/a (short run), vol=0.1%, in_market=100%, exposure=0.01x) 5 decisions A node that calls a model makes the graph live, and nothing else changes. The graph decides five times and trades once, because after the first day it holds the name and answers with an empty action list, which is a decision to change nothing. FinRobot The FinRobot adapter predates the shared layer and is documented in the reference, FinRobot adapter. It needs the finrobot extra on Python 3.11. Decisions and model calls An adapter asks its framework once every every decision steps, and the default is 6. obs.step counts steps across the whole run, so with the default 6 steps a day that is one decision a simulated day, at each day's first step. With steps_per_day=3 the same default asks every second day. On the steps between decisions the adapter records the prices it saw and returns no orders. One decision can cost several calls to the model provider, so a 20-day run is 20 decisions and can be many more provider calls: A framework turn is one model call, and an agent that calls tools takes several turns per decision. The OpenAI Agents adapter caps a decision at max_turns=6, and the PydanticAI adapter at request_limit=8 requests. PydanticAI retries a schema violation inside its own loop, and each retry is another request. The OpenAI Agents SDK 0.22 makes no client-side retry on a malformed answer. A LangGraph graph makes whatever calls its nodes make. The observation payload The framework is shown the serialized payload and never the observation object. The payload is an allowlist written out field by field: Key What it holds step, day, steps_per_day The clock. step counts the whole run. macro The published figures: the VIX, the policy rate, the corporate bond yield, inflation and the business-cycle phase as announced, late. assets Per name: the symbol, price, one-day and five-day returns and volatility from the prices the adapter has seen, the best bid and ask, average daily volume, max_order_shares (the participation cap), the position held, and any fundamentals you pass with fundamentals=. portfolio Cash, net worth, leverage, max_leverage, buying_power and the waiting limit orders. It holds no fair value, no factor attribution and no macro path the run has not reached. That holds at every access level: the adapter's own act receives the read-only market view by default, or the live engine under trusted_agents=True, as What the agent sees sets out, and in both cases the framework receives the payload alone. The decision is a list of actions, each with a symbol, a side (BUY, SELL, HOLD or CANCEL), a positive quantity in shares, and an optional limit_price that makes it a limit order. A bad action is refused on its own and listed in the scorecard's errors. An order over max_order_shares, 2% of the name's average daily volume by default, is cut to the cap, and an order below one share is dropped. Bringing an LLM agent has the full contract. Reference LLM adapters and MCP has every adapter's signature, the shared layer, error types, replay rules and the framework version each adapter was written against. When a run fails or makes no trades, Troubleshooting lists the messages and fixes. ====================================================================== # Local MCP server https://docs.tradefloor.dev/mcp-local.html Run tradefloor's MCP server on your machine so an MCP client can build markets, score strategy specs and explain price moves. It places no orders. ====================================================================== GUIDES/ LOCAL MCP SERVER Local MCP server tradefloor-mcp is an MCP server that runs on your machine and gives an MCP client, such as Claude Code or Claude Desktop, tradefloor's simulation and evaluation tools. A model uses it to study the simulator from outside: it builds universes and scenarios, scores strategies written as data, and explains why a price moved. No tool places an order, and no market lasts from one call to the next. A model that trades runs through an LLM adapter on your machine, or on the hosted app, in beta. Setup Install the extra, which adds the mcp package, and register the server with your client. The server speaks MCP over stdio, so the client starts it as a subprocess. pip install "tradefloor[mcp]" claude mcp add tradefloor -- tradefloor-mcp The second line is for Claude Code. A client configured with a JSON file, as Claude Desktop is, takes the same command. Give the full path to tradefloor-mcp in the environment where you installed it, because the client does not start in your shell: { "mcpServers": { "tradefloor": { "command": "/path/to/venv/bin/tradefloor-mcp" } } } A first tool call In the client, ask the model to call describe_simulator. It returns what the simulator is, what it is measured to reproduce and what it cannot do, and the model should read it before anything else. A strategy is a JSON StrategySpec, never code, so the next step is to validate one and score it beside the baselines on one market. The same calls from Python, through the MCP client library that the extra installs, show what a client receives. The block needs the extra, and was tested with mcp 2.3.0. import asyncio import json from mcp import ClientSession, StdioServerParameters from mcp.client.stdio import stdio_client SPEC = {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}} async def main(): server = StdioServerParameters(command="tradefloor-mcp") async with stdio_client(server) as (read, write): async with ClientSession(read, write) as session: await session.initialize() tools = await session.list_tools() print(len(tools.tools), "tools") checked = await session.call_tool("validate_strategy", {"spec": SPEC}) print("valid:", json.loads(checked.content[0].text)["ok"]) scored = await session.call_tool( "evaluate_strategies", {"strategies": {"mom": SPEC}, "days": 5}) result = json.loads(scored.content[0].text) print("versus buy-and-hold:", result["versus_buy_and_hold"]["mom"]) asyncio.run(main()) 13 tools valid: True versus buy-and-hold: -18690.03 The strategy lost 18,690 against buy-and-hold on that one market, a 5-day run on seed 7. Every result carries a caveats list that says what it cannot show, such as a run too short for the model's validated scope, and a provenance block with the version, preset, seed and universe, so a model summarizing it sees the limits too. One seed is one sample: rank_strategies runs the same specs across seeds with a paired sign test. The tools Tool What it does describe_simulator What the simulator is, what it is measured to reproduce and what it cannot do check_envelope Whether a question is inside the validated scope, before anything runs validate_strategy Checks a strategy spec and returns its fingerprint, without running it build_universe Builds a universe: generated, concentrated in some sectors, or written by hand build_scenario Composes a scenario and shows what it resolves to list_scenarios The shipped scenarios, the constructors and the targets evaluate_strategies Runs strategies and the five baselines on one market rank_strategies Runs them across seeds, with a paired sign test between each pair run_stress_scenario What a scenario does to each strategy, against the same market without it explain_price_move A price move split into the eleven factors that sum to it explain The random draws behind one day for one company start_job, check_job Slow work in the background, up to 252 days A tool that runs a market returns the same result for the same arguments. A refused call comes back as a result with "ok": false and an error message, so the model can correct it. The scoring tools build a new market for each call and run the strategies in it under the read-only market view an ordinary agent gets, so a strategy cannot read fair value or fork the market. The oracle signal is the one exception, and its row says uses_hidden_state. No tool takes a preset or model coefficients. Every tool runs the default, pt-v20, and selecting another is a library call. Limits Limit Value Days in a direct call 60 Days in a background job 252 Decision steps a day 1 to 22, and days × steps at most 60 × 6 in a direct call Names in a universe 2 to 120 Strategies in one call 8 Seeds in rank_strategies 2 to 12, six by default Background jobs running at once 2 A direct call over 60 days is refused, so pass the same arguments to start_job and poll check_job with the id it returns. Jobs live in the server process. A restart of the server, or of the client that started it, loses them and their results. Reference LLM adapters and MCP in the reference lists the tools and limits beside the LLM adapters, under the module tradefloor.mcp. Installation and extras covers install problems. ====================================================================== # Hosted app https://docs.tradefloor.dev/hosted.html Connect an agent or a bot to hosted markets on app.tradefloor.dev, in beta: MCP at /mcp, the HTTP API at /v1 and an Alpaca-shaped API. Hosted tools place orders. ====================================================================== GUIDES/ HOSTED APP Hosted app The hosted app at app.tradefloor.dev keeps simulated markets, called sessions, on a server. An agent or a bot trades them through an MCP server at /mcp, an HTTP API at /v1, or an API shaped like Alpaca's trading API at /broker/{session_id}. Hosted tools place orders, and time in a session moves only when the caller advances it. To run a model on your own machine instead, use the LLM adapters, or the local MCP server, which places no orders. The hosted app is coming soon, and opens in beta. This page shows how agents and bots will connect to it. Setup Sign in to app.tradefloor.dev with your email address (there is no password) and make an API key on the Keys page, /keys. Send it as Authorization: Bearer tfk_... on every request. MCP clients that support OAuth, such as claude.ai, Claude Desktop, Claude Code and Cursor, can sign in instead of using a key, and the Connect page, /connect, shows the configuration for each client. A key has one of three scopes, chosen when it is made: Scope What it can do full The whole API, including forks, analysis, scoring and choosing a session's seed. Keys themselves are made and revoked only in the browser. trader The trading loop only: observe, orders, advance, close, fills, bars, series, events, its own sessions, describe and usage, and the whole Alpaca-shaped API. It opens a session only from a named template, on a seed it never sees. Use it for an agent under test. read-only Reads everything and changes nothing, for dashboards and reports. A call outside the key's scope answers 403 forbidden and names the scope. The loop Every interface runs the same loop on a session: Open a session once and keep the session_id it returns. Observe the session: the clock, quotes, the account and open orders. Decide what to trade from that observation. Place orders, or cancel orders that are still waiting. Advance time by a step or more, and go back to the second step. Close the session. Its report carries caveats that say what the result does and does not show. A market order fills at the start of the next step, when you advance, at the price the order book gives for its size. Right after it is placed its status is accepted. GET /v1/sessions/{id}/manifest returns the session's manifest, from which the tradefloor package rebuilds its market. MCP The MCP endpoint is https://app.tradefloor.dev/mcp, over streamable HTTP. With Claude Code and a key in TF_KEY: claude mcp add --transport http tradefloor https://app.tradefloor.dev/mcp \ --header "Authorization: Bearer $TF_KEY" Other clients, such as claude.ai, Codex and Cursor, take the same URL, and the Connect page shows the configuration for each. A first session is four tool calls: open_session, observe, place_order with a ticker, a side and a whole-share quantity, and advance. To try it, ask the model to open a session, buy 10 shares of the first ticker, advance one step and observe again. close_session ends the session and returns the report. Every write takes an idempotency_key argument, and the tools besides the loop include fork_session, list_fills, get_bars, get_events, list_presets, describe and get_usage. list_presets names the recommended preset, and every session records its own. HTTP API The same loop over /v1, with curl. A write carries an Idempotency-Key header with a new value for each action: export TF_KEY=tfk_... API=https://app.tradefloor.dev/v1 curl -H "Authorization: Bearer $TF_KEY" $API/describe curl -X POST $API/sessions -H "Authorization: Bearer $TF_KEY" \ -H "Idempotency-Key: $(uuidgen)" -H 'content-type: application/json' \ -d '{"universe_size": 5, "seed": 7}' # the answer carries "session_id"; set S to it curl "$API/sessions/$S/observation?view=compact" -H "Authorization: Bearer $TF_KEY" curl -X POST $API/sessions/$S/orders -H "Authorization: Bearer $TF_KEY" \ -H "Idempotency-Key: $(uuidgen)" -H 'content-type: application/json' \ -d '{"ticker": "AAA", "side": "buy", "quantity": 10}' curl -X POST $API/sessions/$S/advance -H "Authorization: Bearer $TF_KEY" \ -H "Idempotency-Key: $(uuidgen)" -H 'content-type: application/json' \ -d '{"steps": 1}' curl $API/sessions/$S/fills -H "Authorization: Bearer $TF_KEY" POST /v1/sessions/{id}/close ends the session and returns its report. Sending a write again with the same Idempotency-Key returns the first result and does nothing twice, so a request that timed out can be resent safely. Alpaca-shaped API A bot written for alpaca-py trades a hosted session when its base URL is https://app.tradefloor.dev/broker/{session_id} and the API key goes where alpaca-py takes the secret. Sessions are opened on the HTTP API, and the facade trades them. Simulated time moves only when the bot calls /tradefloor/advance, which is where a live bot would sleep. import os from alpaca.trading.client import TradingClient from alpaca.trading.enums import OrderSide, TimeInForce from alpaca.trading.requests import MarketOrderRequest KEY = os.environ["TF_KEY"] url = f"https://app.tradefloor.dev/broker/{os.environ['SESSION_ID']}" trading = TradingClient("tradefloor", KEY, url_override=url) trading.submit_order(MarketOrderRequest(symbol="AAA", qty=10, side=OrderSide.BUY, time_in_force=TimeInForce.DAY)) trading.post("/tradefloor/advance", {"until": "next_open"}) print(trading.get_all_positions()) alpaca-py puts /v2 in front of every path, so the advance call goes to /broker/{session_id}/v2/tradefloor/advance, and the app serves the route there as well as without the /v2. The facade serves the account, clock, calendar, assets, positions, market, limit, stop and stop-limit orders with day or gtc, fill activities, latest trades and quotes, snapshots, bars, and news built from the session's event log. It refuses, with a message naming what it does not serve, streaming, trailing-stop, bracket, OCO and OTO orders, notional and fractional orders, ioc, fok, opg and cls, extended hours, and replacing an order. A bot that depends on any of those needs changes beyond the base URL. Hosted sessions and the library Hosted session Library on your machine Order quantity Whole shares only. A fraction is refused, never rounded. Any finite number of shares, fractions included. When a market order fills At the start of the next step, when you advance. When act returns, against the book at that step. Ticks in a step 30 by default, set when the session opens 65 by default Warm-up history history_days up to 260 history_days up to 2,520 Analysis tools check_envelope, evaluate_strategies, rank_strategies and explain_price_move, with tighter caps a call (20 days, 6 seeds and 4 strategies) The local MCP server's limits, or none through the library Troubleshooting Every error has the same body over HTTP and MCP, {"code", "message", "hint"}, with retry_after in seconds when waiting and resending the same request will work. Code Status What to do unauthorized 401 Send a valid API key, or sign in. forbidden 403 The key's scope does not cover the call. Use a key with a wider scope. invalid_order 400 Fix the ticker, the quantity or the price. A fractional quantity lands here. insufficient_buying_power 403 Send a smaller order or reduce a position first. The hint says how many shares would fit, and an order that lowers exposure is always accepted. rate_limited 429 Wait retry_after seconds and resend. quota_exceeded 429 A daily or capacity limit. The message says when it resets. session_closed 409 The session is read-only now. Open a new one. conflict 409 A race, or a reused idempotency key or client_order_id with a different request. not_found 404 No such session or order for this account. The free plan allows 20 names a session, 5 open sessions, and 2,000 simulated days and 300 compute seconds a day. GET /v1/usage, or the get_usage tool, gives the current limits and what is left today, and does not count as a call. A bot with only a trader key that cannot open a session needs open_from_template, or POST /v1/sessions/from-template. ====================================================================== # API index https://docs.tradefloor.dev/api.html tradefloor's public APIs, grouped by task, with a complete alphabetical index of every supported public symbol and the page that documents it. ====================================================================== API REFERENCE/ API INDEX API index tradefloor's public API, grouped by task. Market > Engine > EngineBatch > Universe > Instrument > Macro > OrderBook > News > NewsImpact > TickResult > MatchResult > PriceLevel > SweepCost > Fill > bonds > Engine.submit Agents and evaluation > Agent > Observation > evaluate > leaderboard > rank > Ranking > AgentRecord > Scorecard > StrategySpec > reference_agents > capture_ratio > Portfolio > baselines.Balanced > run_many Scenarios > Scenario > Intervention > Firing > TARGETS > ScenarioValidationError > run_scenario Counterfactuals > World > Limit > Cancel > agree > compare > Agreement > Comparison > Divergence Reproducibility > Checkpoint > branch > replay > RunManifest Execution > tca.analyse > Execution > flow_impact > FlowImpact Data and integrations > edgar > ArrowStream > tradefloor.gym > integrations > MCP server Model > ModelParams > model_preset > preset_record > Presets The complete index below lists every supported public symbol. Integrations tradefloor.gym TradingEnv, a Gymnasium environment. Extra: pip install "tradefloor[rl]". Reference: Gym environment. tradefloor.edgar Snapshots of SEC filings and Universe.from_edgar. Reference: Engine and data. tradefloor.mcp The MCP server, command tradefloor-mcp. Extra: pip install "tradefloor[mcp]". Tool reference: LLM adapters and MCP. tradefloor.integrations.finrobot Runs a FinRobot agent inside a World, with an allowlist that controls which observations the agent sees. Extra: pip install "tradefloor[finrobot]" (Python 3.11 exactly). Usage: Forks and counterfactuals. Shipped studies: examples/rate-shock/ and examples/integrations/finrobot/. tradefloor.integrations.openai_agents Runs an OpenAI Agents SDK agent against a World. Extra: pip install "tradefloor[openai-agents]". Reference: LLM adapters and MCP. tradefloor.integrations.pydantic_ai Runs a PydanticAI agent against a World. Extra: pip install "tradefloor[pydantic-ai]". Reference: LLM adapters and MCP. tradefloor.integrations.langgraph Runs a LangGraph graph against a World. A LangGraph checkpoint holds the graph's workflow state and none of the market's state. The reference lists what each checkpoint holds. Extra: pip install "tradefloor[langgraph]". Reference: LLM adapters and MCP. tradefloor.integrations Adapters that run an OpenAI Agents SDK agent, a PydanticAI agent, a LangGraph graph or any plain Python function in one evaluation loop. Extras: pip install "tradefloor[openai-agents]", "[pydantic-ai]", "[langgraph]". Usage: LLM adapters. Reference: LLM adapters and MCP. Complete index Every symbol in tradefloor.__all__ , plus the optional submodules that extras install, in alphabetical order. A symbol with no link in the reference column is covered under advanced and compatibility exports below. Symbol Module Category Reference Agent tradefloor.harness evaluation reference AgentRecord tradefloor.ranking evaluation reference agree tradefloor.counterfactual counterfactuals reference Agreement tradefloor.counterfactual counterfactuals reference apply_mispricing tradefloor._core model utility exported, with no page of its own ArrowStream tradefloor._core data reference atlas tradefloor.atlas model utility exported, with no page of its own baselines tradefloor.baselines evaluation reference Battery tradefloor.fingerprint evaluation reference battery tradefloor.fingerprint evaluation reference BATTERY_VERSION tradefloor.fingerprint evaluation reference bonds tradefloor market reference boundary tradefloor.boundary counterfactuals reference BoundaryMap tradefloor.boundary counterfactuals reference branch tradefloor.checkpoint reproducibility reference callable tradefloor.integrations.callable integration reference Cancel tradefloor.portfolio execution reference capture_ratio tradefloor.baselines evaluation reference capture_withheld tradefloor.baselines evaluation reference Cell tradefloor.fingerprint evaluation reference characteristic_root_moduli tradefloor._core model utility exported, with no page of its own check_rate tradefloor._core model utility exported, with no page of its own Checkpoint tradefloor.checkpoint reproducibility reference commit tradefloor.fingerprint trust reference common tradefloor.integrations.common integration reference compare tradefloor.counterfactual counterfactuals reference Comparison tradefloor.counterfactual counterfactuals reference counterfactual tradefloor.counterfactual counterfactuals reference crisis_epicentre_solve tradefloor._core model utility exported, with no page of its own crowd_adjusted_root_moduli tradefloor._core model utility exported, with no page of its own DayLedger tradefloor.manifest reproducibility reference Divergence tradefloor.counterfactual counterfactuals reference edgar tradefloor.edgar data reference Engine tradefloor._core market reference EngineBatch tradefloor._core market reference envelope tradefloor.envelope trust reference evaluate tradefloor.harness evaluation reference Execution tradefloor.tca execution reference explain tradefloor.explain trust reference Explanation tradefloor.explain trust reference externalities tradefloor.externality evaluation exported, with no page of its own Externality tradefloor.externality evaluation exported, with no page of its own externality tradefloor.externality evaluation exported, with no page of its own facts tradefloor.facts trust exported, with no page of its own fair_value tradefloor._core model utility exported, with no page of its own FairValue tradefloor._core model exported, with no page of its own Fill tradefloor._core market reference Fingerprint tradefloor.fingerprint evaluation reference fingerprint tradefloor.fingerprint evaluation reference FingerprintComparison tradefloor.fingerprint evaluation reference Firing tradefloor.interventions scenarios reference Flip tradefloor.boundary counterfactuals reference flip tradefloor.boundary counterfactuals reference flow_impact tradefloor execution reference FlowImpact tradefloor execution reference GameRng tradefloor._core utility exported, with no page of its own gym tradefloor.gym integration reference HiddenState tradefloor.sandbox evaluation reference History tradefloor.harness evaluation reference impulse_response tradefloor._core model utility exported, with no page of its own Instrument tradefloor._core market reference integrations tradefloor.integrations integration reference Intervention tradefloor.interventions scenarios reference interventions tradefloor.interventions scenarios reference Invariance tradefloor.counterfactual integration reference invariance tradefloor.counterfactual integration reference JSONRenderer tradefloor.render integration reference langgraph tradefloor.integrations.langgraph integration reference leaderboard tradefloor.harness evaluation reference Limit tradefloor.portfolio execution reference loss tradefloor.loss trust exported, with no page of its own Macro tradefloor._core market reference manifest tradefloor.manifest reproducibility reference map_boundaries tradefloor.boundary counterfactuals reference market_status tradefloor._core utility exported, with no page of its own MarketView tradefloor.sandbox evaluation reference MatchResult tradefloor._core market reference mcp tradefloor.mcp integration reference MispricingState tradefloor._core model exported, with no page of its own model_preset tradefloor._core model reference ModelParams tradefloor._core model reference News tradefloor._core market reference NewsImpact tradefloor._core market reference Node tradefloor.explain trust reference noise tradefloor.noise reproducibility exported, with no page of its own Observation tradefloor.harness evaluation reference openai_agents tradefloor.integrations.openai_agents integration reference oracle_is_ceiling tradefloor.baselines evaluation reference OrderBook tradefloor._core market reference OrderError tradefloor._core errors reference Portfolio tradefloor.portfolio evaluation reference PortfolioView tradefloor.sandbox evaluation reference Position tradefloor.portfolio evaluation exported, with no page of its own preset_names tradefloor._core model reference preset_record tradefloor.records model reference preset_records tradefloor.records model reference PriceLevel tradefloor._core market reference pydantic_ai tradefloor.integrations.pydantic_ai integration reference rank tradefloor.ranking evaluation reference Ranking tradefloor.ranking evaluation reference RATE_SECTOR tradefloor market reference rate_specs tradefloor market reference RATE_TICKERS tradefloor market reference reference_agents tradefloor.baselines evaluation reference render tradefloor.render integration reference Renderer tradefloor.render integration reference replay tradefloor.replay reproducibility reference Resample tradefloor.counterfactual counterfactuals reference resample tradefloor.counterfactual counterfactuals reference reveal tradefloor.fingerprint trust reference run_many tradefloor market reference run_scenario tradefloor.scenario scenarios reference RunManifest tradefloor.manifest reproducibility reference sandbox tradefloor.sandbox evaluation reference SandboxError tradefloor.sandbox errors reference Scenario tradefloor.scenario scenarios reference ScenarioValidationError tradefloor.interventions errors reference Scorecard tradefloor.harness evaluation reference sealed_battery tradefloor.fingerprint evaluation reference sector_daily_sigma tradefloor._core model utility exported, with no page of its own sectors tradefloor._core utility exported, with no page of its own spec tradefloor.spec evaluation reference SPEC_VERSION tradefloor.spec evaluation exported, with no page of its own stationary_sigma tradefloor._core model utility exported, with no page of its own step_mispricing_daily tradefloor._core model utility exported, with no page of its own StrategySpec tradefloor.spec evaluation reference sweep tradefloor.sweep evaluation exported, with no page of its own SweepCost tradefloor._core market reference TARGETS tradefloor.boundary scenarios reference tca tradefloor.tca execution reference TextRenderer tradefloor.render integration reference TickResult tradefloor._core market reference TradingEnv tradefloor.gym integration reference Universe tradefloor market reference UNSUPPORTED_TARGETS tradefloor.interventions scenarios reference ValidationError tradefloor._core errors reference Verification tradefloor.manifest reproducibility reference version tradefloor._core utility exported, with no page of its own __version__ tradefloor utility exported, with no page of its own versus_buy_and_hold tradefloor.baselines evaluation reference World tradefloor.counterfactual counterfactuals reference yaml_subset tradefloor.yaml_subset utility exported, with no page of its own Advanced and compatibility exports Every symbol below is in tradefloor.__all__ and stays there for the life of the release. For a symbol with no page of its own, this table is the reference. It gives the import path, the signature and how far you can rely on the symbol. supported for normal use advanced a full API that few callers need compatibility kept so older code keeps working legacy may be removed in a later release Signature Import Stability Purpose Portfolio(cash: 'float' = 1000000.0, *, max_leverage: 'float | None' = None, cash_interest: 'bool' = False, owner: 'str' = 'agent', margin_interest: 'bool' = True) -> 'None' tradefloor.portfolio supported Cash, positions and P&L for one trader. Members: accrue, cancel, cash, cash_interest, clear_flow, execute, fills, fills_table, gross_exposure, interest, leverage, margin_interest, market_value, marks, max_leverage, net_worth, open_orders, owner, pending_flow, pnl, positions, quantity_of, realised, stamp, starting_cash, submit_limit, sync, unrealised. Position(ticker: 'str') -> 'None' tradefloor.portfolio supported A holding in one instrument. Members: avg_cost, market_value, quantity, realised, ticker, unrealised. SPEC_VERSION: int = 1 tradefloor supported The StrategySpec schema version. A spec written against an older number still loads. One written against a newer number does not, so a manifest that records it says whether this install can replay the run. sweep(seeds: 'Iterable[int]', *, universe: 'Sequence[Instrument]', days: 'int' = 1, ticks_per_day: 'int' = 390, macro: 'Macro | None' = None, scenario: 'Any' = None, collect: 'str' = 'bars', grain: 'str | None' = None, minutes: 'int | None' = None, start: 'tuple[int, int, int]' = (9, 30, 3), workers: 'int' = 1, model: 'str | ModelParams | None' = None) -> 'Iterator[tuple[int, Any]]' tradefloor.sweep supported Yield (seed, table) for each seed, one at a time. __version__: str = '0.8.7' tradefloor compatibility The installed package version, as a string. It is what a RunManifest records and what a published result cites. apply_mispricing(fair_value, s) tradefloor._core advanced Price from fair value and mispricing: fair_value * exp(s), floored. tradefloor.atlas tradefloor.atlas advanced Atlas: a map of how the model responds, instead of a search that guesses. characteristic_root_moduli(phi=None, theta=None) tradefloor._core advanced Moduli of the AR(2) characteristic roots. check_rate(name, fraction) tradefloor._core advanced Validate a fractional rate, raising ValueError if it looks wrong. crowd_adjusted_root_moduli() tradefloor._core advanced The same, with the crowd-feedback term included. tradefloor.facts tradefloor.facts advanced Stylised facts: what these markets look like, measured, next to real ones. fair_value(*, eps=None, sector, revenue_growth=None, federal_funds_rate=0.0, corporate_bond_yield=None, qe_pe_boost=None, book_value_per_share=None, neutral_discount_rate=None) tradefloor._core advanced Fair value per share from fundamentals and the macro state. FairValue tradefloor._core advanced A fair-value result and its decomposition. Members: book_value_path, fair_value, qe_adjustment, rate_adjustment, sector_anchor_pe, target_pe. GameRng(seed, sequence) tradefloor._core advanced Deterministic seeded random number generator (PCG32 + Box-Muller). Members: next_bool, next_float, next_int, next_normal. impulse_response(horizon_days, phi=None, theta=None) tradefloor._core advanced Impulse response of the daily process over horizon_days. tradefloor.loss tradefloor.loss advanced The calibration objectives: band distance, and a residual scoring rule. market_status(hour, minute, day_of_week) tradefloor._core advanced Which session a moment falls in. MispricingState(s=0.0, s_prev=None) tradefloor._core advanced The daily log-mispricing state. Members: s, s_prev. sector_daily_sigma(sector) tradefloor._core advanced A sector's long-run daily return standard deviation, as a fraction. sectors() tradefloor._core advanced The twelve sector keys, in contractual declaration order. stationary_sigma(innovation_sigma, *, phi=None, theta=None) tradefloor._core advanced Standard deviation of the daily mispricing process at rest. step_mispricing_daily(state, *, innovation=0.0, shock=0.0) tradefloor._core advanced One DAILY step of the mispricing process. version() tradefloor._core compatibility The compiled core's version, as a string. It reports the same value as __version__ and is kept for callers written before that attribute existed. tradefloor.yaml_subset tradefloor.yaml_subset advanced A reader for the boring part of YAML, and nothing else. Atlas surveys the model's parameter space. It reports which parameters move which outputs, what shape each effect has, and which trades are available at once. It is meant for people who calibrate the model, and these docs have no tutorial for it. ====================================================================== # Engine and data https://docs.tradefloor.dev/core-types.html Run a market and read what it made: Engine, Universe and Macro, the Arrow output tables, real companies from EDGAR, units, and every typed member. ====================================================================== API REFERENCE/ ENGINE AND DATA Engine and data Build a market, run it, and read back what it made. Running a market import tradefloor as tf universe = tf.Universe.random(20, seed=11) macro = tf.Macro(federal_funds_rate=0.025, corporate_bond_yield=0.052, vix=16.0) engine = tf.Engine(seed=42, universe=universe, macro_state=macro) engine.run_days(60) bars = engine.bars(grain="day") # OHLCV per name per day truth = engine.truth() # fair value, mispricing, eleven factors economy = engine.macro_table() # the macro state, one row per day Three inputs fix a market. The universe is the roster of companies, in order: the engine draws random numbers in roster order, so a re-sorted roster is a different market from the same seed, and universe.fingerprint covers order as well as content. The Macro is the economy on day zero only, because the engine moves rates, inflation, the business cycle and the VIX forward at every close after that. The seed fixes every random draw. There are two seeds. The one passed to Universe.random picks each company's fundamentals, and the one passed to Engine picks the market. To find how much of a result was luck, hold the universe fixed and change the engine's seed. A seed can be any integer from 0 to 2**64 - 1. A pin replaces the model's own update of a macro field once, and the next days move on from it. To hold a field on every day, pin it every day or use a scenario. Output tables Call What it holds bars(grain="day") OHLCV per name per bar. Volume is the shares traded in the bar. truth() Fair value, mispricing and the eleven factor contributions behind every move, per tick. macro_table() The economy, one row per day. take_fills() What an agent asked for against what it got. book_table() Recorded book depth, per tick, side and level. Each is an Arrow table, which polars, pandas, pyarrow and duckdb read without copying, and tradefloor depends on none of them. Value columns are float64 throughout, and a missing value is NaN, never zero. Universe.random(20, seed=101, bonds=True) adds three rate indices, UST2Y, UST10Y and IGCORP, priced from the simulated yield curve, so a portfolio can hold bonds beside the stocks. Real companies from EDGAR snap = tf.edgar.fetch(as_of="2024-06-30", limit=100, user_agent="Jane Roe jane@example.org") snap.save("edgar-2024h1.json") # hashed, so it can be cited universe = tf.Universe.from_edgar(snap, federal_funds_rate=0.03) The universe keeps the spread of valuations, the sector weights and the share of loss-making companies in the filings. Its price paths are still simulated. user_agent must carry a contact address, as the SEC's fair-access policy asks. EDGAR can revise past records, so cite the saved snapshot file. rank_by="equity", the default, favors large balance sheets; rank_by="public_float" gives a list closer to a real index. Units and conventions The engine does not clamp what you pass it: a malformed input raises ValidationError and a refused order raises OrderError. Some layers above it resize orders. An LLM adapter cuts an order to the participation cap and records the cut, tf.baselines.rebalance caps each trade at a share of the name's average daily volume, 2% by default, and the gym environment clips an action to [-1, 1] and scales one over the leverage cap down, setting info["scaled"]. Write Not Why 0.052 5.2 Rates are decimals. 5.2 raises an error that says so. None 0.0 corporate_bond_yield=None falls back to the policy rate; 0.0 is a real value. eps=-1.20 filtering losses out A loss-making company is valued from book value, a path the model needs. 3_000_000 0.03 Short interest is a number of shares. model="pt-v20" Engine(garch_alpha=...) Coefficients are chosen by preset name, so two results can be compared. Type aliases Literal types the signatures below refer to. Name Accepted values ColumnField Literal[ "price", "previous_close", "previous_tick_price", "open", "high", "low", "volume", "avg_volume", "market_cap", "mispricing_s", "mispricing_s_prev_close", "mispricing_momentum", "last_daily_return", "maker_inventory", "garch_variance", "beta", "short_interest", "float_shares", ] CycleName Literal["expansion", "peak", "contraction", "trough", "recovery"] FactorName Literal[ "reversion", "momentum", "crowd_lean", "company_news", "order_flow_impact", "short_squeeze_effect", "random_noise", "circuit_breaker", "jump", "overnight", "fair_value_shift", ] Grain Literal["tick", "day"] MarketStatusName Literal["open", "pre_market", "after_hours", "closed"] Side Literal["buy", "sell"] Engine A running market. You build it from a seed (the number that fixes every random draw), a roster and the economy on day zero, then advance it a day or a tick at a time. run_session changed in 0.8.5: an agent's trades go in fills= and a standing rate in flow_per_tick=, and order_flow= raises ValidationError. The 0.8.5 release notes say why and what it changes. Engine( *, seed: int, universe: Sequence[Instrument], macro_state: Macro | None = None, model: str | ModelParams | None = None, ) -> None Running the market Signature Meaning def run_days( days: int, *, hour: int = 9, minute: int = 30, day_of_week: int = 3, ticks_per_day: int = 390, volatility: float = 1.0, record: bool = True, first_day: int | None = None, ledger: Any | None = None, ) -> int Advance whole trading days: open, session, close, repeat. The usual way to run a market. def run_session( hour: int, minute: int, day_of_week: int, ticks: int, *, volatility: float = 1.0, close_at_end: bool = False, news: Sequence[News] | None = None, news_impacts: Sequence[NewsImpact] | None = None, fills: dict[str, tuple[float, float]] | None = None, flow_per_tick: dict[str, tuple[float, float]] | None = None, order_flow: None = None, ) -> int Run many ticks inside one day, between an explicit open and close. fills takes an agent's trades as {ticker: (bought, sold)} in shares and applies them once, on the session's first tick. flow_per_tick takes the same shape and applies it on every tick, as a standing rate. Since 0.8.5 order_flow raises ValidationError, and the note at the top of this section links to the release notes. def tick( hour: int, minute: int, day_of_week: int, *, volatility: float = 1.0, news: Sequence[News] | None = None, news_impacts: Sequence[NewsImpact] | None = None, order_flow: dict[str, tuple[float, float]] | None = None, ) -> TickResult Advance one game-minute. order_flow is what traders bought and sold in that minute, as {ticker: (bought, sold)} in shares, and 0.8.5 left it unchanged. def run_until( *, ticker: str, above: float | None = None, below: float | None = None, max_ticks: int = 390, hour: int = 9, minute: int = 30, day_of_week: int = 3, volatility: float = 1.0, ) -> int | None Advance until a named instrument's price leaves a band, or until max_ticks elapses. Refuses a rate index, whose level moves only when a yield is written, between sessions or by pin_macro. def open_market(*, day: int | None = None) -> None Roll the day's opening marks and draw the day's endogenous news. Call once before a session's ticks. def close_market() -> None Run the close bookkeeping: the variance update, the daily roll and the end-of-day state. session_tick: int | None Ticks run since the current day's open. It reads 0 at the open, 390 after a full session, and stays at 390 after the close until the next open. None until this engine, or one restored from a snapshot, has opened a day. It reads a counter the engine already keeps, so it cannot change a run. Reading what happened Signature Meaning def bars( *, day: int | None = None, minutes: int | None = None, grain: Grain | None = None, ) -> ArrowStream OHLCV per instrument, as an Arrow stream. A bar's volume is the shares traded inside it, at every grain, so a day's tick rows add up to its day bar. day=None (the default) returns every recorded day. day=N returns that day alone, and an unrecorded day raises. Output tables lists the columns. def truth(*, day: int | None = None) -> ArrowStream The labeled dataset: true value, mispricing and the eleven factor contributions behind every simulated move. day=None (the default) returns every recorded day. An unrecorded day raises. def macro_table() -> ArrowStream One row per recorded day of the macro state. Row d holds the values day d traded under, recorded before its close, so it is what day d - 1's close produced. To pair a day's return with the macro move it caused, compare row d + 1 with row d. The values after the last recorded close are on macro_state, not in the table. def book_table() -> ArrowStream Recorded book depth, one row per tick, instrument, side and level. def prices() -> bytes Current price per instrument, as little-endian f64 bytes in roster order. def column(field: ColumnField) -> bytes One named column across every instrument, as little-endian f64 bytes in roster order. def attribution(factor: FactorName) -> bytes One named factor's contribution across every instrument, as little-endian f64 bytes. Zero in every rate index's slot, because an index moves by the repricing formula that rate_attribution splits. def session_prices() -> bytes The last session's price path, ticks_written by instruments, row-major. def session_volumes() -> bytes The last session's volume path, same shape as session_prices. Each value is the name's running total for the day, which the open resets to zero. def session_mispricing_s() -> bytes The last session's mispricing path, same shape as session_prices. def session_news() -> list[dict[str, Any]] The current news day's endogenous events, one dict each with ticker, sector, price_impact and day. ticker is None when the event names no company on the roster. The list is empty before the first open and on a preset with endogenous_news_intensity at 0. price_impact is the whole move the event adds by the close, which makes it the answer key, so never pass it to an agent. It is a read. It takes no draw and cannot change a run. def noise_split( part: Literal["market", "sector", "idio"], ) -> bytes The day's random_noise column split into the three draws it sums, for part='market', 'sector' or 'idio', as little-endian f64 bytes in roster order. It is a read. It takes no draw and cannot change a run. The order book Signature Meaning def book(ticker: str) -> OrderBook A detached OrderBook for one instrument, as it stands now. def snapshot_book( *, day: int = 0, tick: int = 0, levels: int = 10, ) -> int Record current depth for every instrument into the book table. Orders from agents Signature Meaning def submit( agent: str, ticker: str, quantity: float, *, limit_price: float | None = None, order_id: str | None = None, ) -> dict[str, Any] Send one agent's order to this engine's book and return a report dict. quantity is signed, positive to buy. limit_price=None is a market order, and what the book cannot fill comes back as unfilled. A limit order's unfilled part waits in the book's queue with book_resting on. With it off, the part waits outside the book and fills in full at its limit on the first tick whose print reaches it. Every share taken reaches the market once, on the next tick the market is open. The report carries requested, filled, average_price, worst_price, reference (the mid the order met), resting, unfilled, mode and fills. Raises OrderError for a rate index and for the labels mm, depth, flow and range. Recorded in the order log. def submit_many( orders: Sequence[dict[str, Any]], ) -> list[dict[str, Any]] Send several agents' orders for one step in a fixed order. Each dict has agent, ticker and quantity, and optionally limit_price and order_id. Orders run sorted by agent label, and one agent's in list order, so the same orders give the same market however the list was built. Each order meets the book the orders before it left. Returns one report per order, in the order processed. A refused order raises, and the orders before it stand. def cancel( order_id: str, *, agent: str | None = None, ) -> bool Cancel a waiting order by id, and with agent only if that agent sent it. Returns whether an order was removed. Recorded in the order log. def open_orders( agent: str | None = None, ) -> list[dict[str, Any]] Waiting orders for one agent, or for every agent with None, in arrival order. Each dict has order_id, agent, ticker, side, limit_price, quantity, remaining, sequence and mode, which is queue or range. Changes nothing. def take_fills( agent: str | None = None, ) -> list[dict[str, Any]] Collect the fills the book holds for one agent, or for all, in the order they happened, and forget them. liquidity is taker, maker or range, and counterparty is mm, depth, flow, range or another agent's label. A waiting order that the market's own flow filled during a session is reported here and nowhere else. Recorded in the order log. def take_impacts( agent: str | None = None, ) -> list[dict[str, Any]] Collect the permanent impact of each agent's flow, one row per agent, name and tick, and forget it. permanent is the change the flow made to the name's mispricing_s, in log units. It is exact under fill_impact_coefficient, and otherwise the tick's order-flow impact shared among agents by signed shares. Recorded in the order log. def book_live() -> bool True when book_shared or book_resting is on, so an agent's order executes in this engine's book and Portfolio.execute sends it there. True on pt-v20 and False on every earlier preset. Rate indices Signature Meaning def rate_instruments() -> list[dict[str, Any]] One dict per rate index this engine holds, in roster order, with ticker, name, curve_point, duration, convexity, spread_bps, level, yield, price and avg_volume. Empty on an engine without them. def curve() -> dict[str, float] The curve now, as fractions: policy_rate, treasury_2y, treasury_10y and corporate, the yield equities are discounted off. investment_grade, the yield IGCORP reads, is there only on an engine holding rate indices. def rate_attribution( component: Literal["carry", "duration", "convexity", "flow"], ) -> bytes One rate component of today's move for every instrument, as little-endian f64 bytes in roster order, zero for each equity. carry, duration and convexity are the terms of the repricing formula summed over the day, as fractions of the level. flow is the print's premium over the index level now, price / level - 1. Reset at each open. RATE_COMPONENTS: list[str] The components rate_attribution takes, in order: carry, duration, convexity and flow. Recording Signature Meaning def record(day: int) -> None Capture the session just run, and the macro state, as one recorded day. def clear_recording() -> None Discard every recorded day. The market itself is untouched. recorded_days: int How many days have been recorded. recorded_book_rows: int How many book rows have been recorded. session_ticks_written: int Ticks written by the last session. Market state Signature Meaning def fork(count: int = 2) -> list["Engine"] Copy the engine, including the order book, the day's endogenous news and the generator position. Returns count independent engines identical to this one. tf.branch calls it. Orders waiting in the agent-facing book are copied too. def state_snapshot() -> dict[str, Any] Market state as one dict: every column, the variance and crisis state, and the generator position on every stream. It gains a rates entry on an engine holding rate indices, and a book entry once an agent has sent an order. def state_hash() -> str This market's state as one 64-character hex digest: the ledger leaf, covering every field state_snapshot carries. manifest.state_hash computes the same digest in Python. The field set grew at 0.8.0, so a state hashes differently under 0.7.x and 0.8.0 even where the prices agree. From 0.8.5 it also covers the rate indices and the agent-facing book where the engine has them, so an engine with neither hashes as it did under 0.8.1. def restore_state(snapshot: dict[str, Any]) -> None Put a market back to a captured state. A snapshot taken under 0.7.x is refused, because the engine now carries ten random streams where 0.7.x carried eight. order_log: list[dict[str, Any]] Every input that crossed into this engine, in order, including every submit, cancel, take_fills and take_impacts call. What a Checkpoint replays. A run_session entry names its flows fills and flow_per_tick, and a log written by 0.8.x names one order_flow, which replays as flow_per_tick. draws_consumed: int Cumulative draws across all three streams. Two runs that agree here consumed the same randomness. def draws_by_stream() -> dict[str, int] Cumulative draws on three of the streams, as {'market': n, 'economy': n, 'external': n}. stream_positions reports all ten. def stream_positions() -> dict[str, tuple[int, int]] (uniforms, normals) taken so far on each stream, keyed by stream name: the address the next draw of each kind would take. def draw_normal() -> float One normal draw from the engine's external stream. def draw_uniform() -> float One uniform draw from the engine's external stream. The roster Signature Meaning tickers: list[str] Instrument tickers, in roster order. len: int The number of instruments on the roster. def index_of(ticker: str) -> int | None The roster index of a ticker, or None. def list_instrument(instrument: Instrument) -> int Add an equity to a running market, after the last equity. Returns its index. A rate index is refused, because rate indices are fixed when the engine is built. def delist(index: int) -> str Remove the instrument at an index, returning its ticker. A rate index cannot be delisted. Model and economy Signature Meaning model: ModelParams The coefficient set this engine runs. model_params: dict[str, Any] The full coefficient dictionary, as ModelParams.to_dict() returns it. model_fingerprint: str The model fingerprint: the preset name for an unmodified preset, or custom- after any override. macro_state: Macro The current macro state. macro_fields: dict[str, Any] Every field pin_macro writes, in the units it takes. Distinct from macro_state, the Macro object. def pin_macro( *, vix: float | None = None, federal_funds_rate: float | None = None, corporate_bond_yield: float | None = None, inflation_rate: float | None = None, qe_pe_boost: float | None = None, qe_assets_ratio: float | None = None, fear_greed_index: float | None = None, gdp_growth: float | None = None, unemployment_rate: float | None = None, tariff_rate: float | None = None, oil_price: float | None = None, cycle: CycleName | None = None, epicentre: str | None = None, vix_sets_variance: bool = False, treasury_yield_2y: float | None = None, treasury_yield_10y: float | None = None, ) -> None Hold one or more macro series at given values instead of letting them evolve. epicentre names the sector that carries the next crisis episode, or 'none', and persists once written. vix_sets_variance=True makes tonight's close set the market's variance from the VIX pinned today, where it would otherwise move one step toward it. Everything is validated before anything is written. treasury_yield_2y and treasury_yield_10y write the curve the rate indices read, as fractions. The chain carries on from a pinned yield, and every close recomputes the 2-year from the policy rate and the 10-year, so a 2-year pinned alone lasts until that close. def vix_sets_variance_pending() -> bool True when tonight's close will set the market factor's variance from the VIX, because the VIX was pinned today with vix_sets_variance=True. The close clears it. def crisis_episode() -> tuple[bool, int, str | None] The crisis episode, as (in_episode, sessions_under, epicentre). An episode starts when the VIX crosses crisis_vix_threshold and ends after crisis_epicentre_end_sessions consecutive sessions under it, and sessions_under counts toward that end. epicentre is a sector key, 'none' for a crisis with no epicentre, or None when no episode is running. It is always (False, 0, None) on a preset with crisis_epicentre_extra at 0.0, which is every preset before pt-v19. def set_avg_volume(values: Sequence[float]) -> None Write the avg_volume column the market maker quotes off, one value per instrument in roster order. FACTORS: list[str] The eleven factor names, in the order truth() reports them. An Engine is mutable. Every run and tick method advances it in place and returns None unless the table says otherwise. fork() returns independent copies of the current engine state. Checkpoint is the serialized form, which you can save and load in a later process. From 0.8.0, state_snapshot carries thirteen more fields, including the crisis episode, the sector variance and the slow variance levels, and state_hash covers them. So the same prices hash differently under 0.7.x and 0.8.0. A snapshot taken under 0.7.x cannot be restored, because the engine now draws from ten random streams where 0.7.x drew from eight. From 0.8.5, state_snapshot and state_hash also cover the rate indices on an engine that holds them, and the agent-facing book once an agent has sent it an order. An engine with neither hashes as it did under 0.8.1. Instrument One tradable company. Every monetary field is in currency units per share unless it says otherwise. Instrument( ticker: str, sector: str, *, initial_price: float, shares_outstanding: float, eps: float | None = None, book_value_per_share: float | None = None, revenue_growth: float | None = None, avg_volume: float = 1000000.0, beta: float = 1.0, short_interest: float = 0.0, ) -> None Fields Signature Meaning ticker: str The instrument's symbol. Unique within a universe. sector: str One of the twelve sector names. initial_price: float Day-zero price, in currency units per share. shares_outstanding: float Share count. eps: float | None Earnings per share, in currency units. None marks a loss-maker, which is priced on book value instead. book_value_per_share: float | None Book value per share, in currency units. revenue_growth: float | None Revenue growth as a decimal fraction: 0.08 is 8%. avg_volume: float Average daily volume, in shares. beta: float Loading on the market factor. short_interest: float Short interest as a SHARE COUNT, not a fraction of float. market_cap: float Derived, read-only: initial_price times shares_outstanding. A simulated rate index is an Instrument too, with sector "rates", and tf.bonds() builds all three. Instrument() accepts that sector only for UST2Y, UST10Y and IGCORP, and refuses eps, book_value_per_share, revenue_growth or a non-zero short_interest on one. Macro The economy on day zero. Rates are always decimal fractions, so 0.052 is 5.2%. Passing 5.2 raises ValidationError, and the message names the fraction you probably meant. Macro( *, vix: float = 15.0, federal_funds_rate: float = 0.025, corporate_bond_yield: float | None = None, inflation_rate: float = 0.02, qe_pe_boost: float = 0.0, qe_assets_ratio: float = 1.0, fear_greed_index: float = 50.0, cycle: CycleName = 'expansion', ) -> None Fields Signature Meaning vix: float Volatility index level, in points. federal_funds_rate: float Policy rate, as a decimal fraction. corporate_bond_yield: float | None Corporate yield, as a decimal fraction. None falls through to the policy rate plus a spread. inflation_rate: float Inflation, as a decimal fraction. qe_pe_boost: float Additive boost to the target price/earnings multiple. qe_assets_ratio: float Central-bank asset stock relative to the neutral portfolio. 1.0 is neutral. Enters fair value as a concave stock term. fear_greed_index: float Sentiment index, 0 to 100. cycle: str The business-cycle phase. This is the state on day zero only, and every close advances it. To set a path for the whole run, use a Scenario, or hold fields still with Engine.pin_macro. OrderBook One company's order book. Orders match against resting orders on price-time priority (best price first, then earliest), so an order's price depends on how many levels it uses up. OrderBook( company_id: str, last_price: float | None = None, ) -> None Members Signature Meaning company_id: str The instrument this book belongs to. best_bid: float | None Best resting bid, or None when that side is empty. best_ask: float | None Best resting ask, or None when that side is empty. mid_price: float | None Midpoint of the touch, or None when either side is empty. spread: float | None Ask minus bid, or None when either side is empty. def post_limit( side: Side, price: float, quantity: float, *, owner: str, order_id: str | None = None, ) -> str Rest a limit order on the book. Returns its id. def submit( side: Side, quantity: float, *, taker: str = 'taker', limit_price: float | None = None, post_remainder: bool = False, order_id: str | None = None, ) -> MatchResult Submit an order against the book. Matches against resting depth. def append_maker_level( side: Side, price: float, quantity: float, *, owner: str, ) -> str | None Append a level to the end of one side, skipping the sorted insert. def sweep_cost( side: Side, quantity: float, ) -> SweepCost | None What sweeping a quantity would cost, without executing it. None when the side cannot fill it. def price_levels( side: Side, max_levels: int = 10, ) -> list[PriceLevel] Aggregated levels on one side, best first. def depth(side: Side) -> float Total resting quantity on one side, in shares. def cancel_order(order_id: str) -> bool Cancel one resting order by id. True when it was there. def cancel_all_for(owner_id: str) -> int Cancel every resting order owned by one party. Returns how many. A book from Engine.book() is a detached copy. Reading it or filling orders against it leaves the engine unchanged, byte for byte. PriceLevel One price level with its resting orders added together, as price_levels() returns it. PriceLevel Fields Signature Meaning price: float The level's price. quantity: float Total resting quantity at that price, in shares. orders: int How many separate orders stand at that price. SweepCost What a sweep, one order that takes several price levels, would cost. sweep_cost() returns it without executing anything. SweepCost Fields Signature Meaning average_price: float Volume-weighted average price the sweep would pay. worst_price: float The last, worst price the sweep would reach. filled: float How many shares would fill, which may be fewer than asked. MatchResult The outcome of submit(). MatchResult Fields Signature Meaning fills: list[Fill] The fills the order produced, in match order. unfilled: float Shares that could not fill. Zero when a remainder was posted. average_price: float | None Volume-weighted average across the fills, or None if nothing filled. resting_order_id: str | None Id of the resting remainder, when one was posted. Fill One match between an incoming order and a resting one. Fill Fields Signature Meaning price: float The RESTING order's price, never the incoming one. quantity: float Shares filled. maker_order_id: str Id of the resting order. maker_id: str Owner of the resting order. taker_id: str Party that crossed the spread. taker_side: str Which side the taker was on. Universe The roster, meaning the list of companies in a market, followed by any rate indices. It is a list subclass, so indexing and iteration work as usual. Signature Meaning Universe(instruments: Sequence[Instrument]) Build a roster from instruments you made yourself. Universe.random( n: int = 108, *, seed: int = 0, bonds: bool = False, ) -> Universe Generate a roster, filling twelve sectors in turn. Tickers follow roster position: AAA, AAB, AAC. bonds=True, new in 0.8.5, appends UST2Y, UST10Y and IGCORP after the n equities, which are the same n names either way. Universe.from_edgar(snapshot, **kwargs) -> Universe Build from an EDGAR snapshot, as tf.edgar.fetch() returns it. See Real companies from EDGAR. Universe.from_json(text: str) -> Universe Rebuild from to_json() output. to_json(**kwargs) -> str Serialize the roster to JSON. tickers() -> list[str] Tickers, in roster order. fingerprint: str A sha256 hash of the roster in its standard serialized form, including its order. with_bonds( tickers: Sequence[str] | None = None, ) -> Universe New in 0.8.5. A new roster with rate indices appended, the ones named in the order given or all three. Leaves this roster alone, and refuses one that already holds any of them. equities() -> Universe New in 0.8.5. A new roster of the equities alone, in roster order. Roster order is part of the universe's identity and is included in fingerprint . If you reorder the roster, the same seed gives a different market. Reproducibility explains how the fingerprint is built. Rate indices Three simulated rate indices can trade beside the equities from 0.8.5. UST2Y and UST10Y are constant-maturity 2-year and 10-year treasury indices, and IGCORP is an investment-grade corporate bond index. None is a real security. They are priced off the engine's own curve and take no random draws, so every equity price, draw and macro value is identical to the same run without them. Signature Meaning tf.bonds( tickers: Sequence[str] | None = None, ) -> list[Instrument] The rate indices as instruments with their default terms, the ones named or all three. An unknown ticker raises ValidationError. tf.rate_specs() -> list[dict[str, Any]] Each index's fixed terms: ticker, name, curve_point, duration, convexity, spread_bps, avg_volume, units_outstanding and initial_price. tf.RATE_TICKERS: tuple[str, ...] The three tickers in their default order: UST2Y, UST10Y and IGCORP. tf.RATE_SECTOR: str The sector every rate index carries, "rates". tf.sectors() does not list it. Each index reads one yield. UST2Y reads the 2-year treasury yield and UST10Y the 10-year. IGCORP reads the 10-year plus a credit spread, which is re-marked whenever the engine sets the corporate yield or a caller pins it. When that yield changes by dy , the level moves by carry - D * dy + 0.5 * C * dy**2 , where carry is the yield over 252 at the first open after a close and zero otherwise. Duration D and convexity C are 1.9 and 4.6 for UST2Y, 8.5 and 84 for UST10Y, and 7.0 and 100 for IGCORP. The economy moves yields only at the close, so a level holds still through a session unless pin_macro writes a yield. The level reprices at the next open, or on the first tick after the pin. Nothing interpolates toward a later yield, so no price shows a yield before the close that sets it. Engine.rate_attribution splits the day's move into carry, duration, convexity and the book's flow. An index trades like an equity through Portfolio.execute , the fills= argument of run_session , the tape and tca.analyse , on a book of its own. That book is the maker's ladder around the index level, 0.6 basis points wide before cent rounding for the treasuries and 0.8 for IGCORP, and it widens with the VIX by the equity rule. Its depth comes from the median dollar volume of SHY, IEF and LQD over 2015 to 2025. The maker lays off a trade's inventory with a 15-minute half-life, so a trade's impact on an index fades. Rate indices come after every equity. An engine refuses a roster that lists an equity after one, lists one twice, or holds rate indices and no equities. The agent-facing book holds equities only, so Engine.submit and Portfolio.submit_limit refuse a rate index. So do run_until , a News item that names one, explain , list_instrument , delist and EngineBatch . To move an index, write the yield it reads with pin_macro(treasury_yield_2y=..., treasury_yield_10y=...) , or with a scenario on macro.treasury_2y or macro.treasury_10y . import tradefloor as tf u = tf.Universe.random(20, seed=101, bonds=True) u.tickers()[-3:] # ['UST2Y', 'UST10Y', 'IGCORP'] e = tf.Engine(seed=3, universe=u) e.run_days(5) e.curve["treasury_10y"] e.rate_instruments[1]["level"] On pt-v20, the default, the 2-year yield moves 3.87 basis points a day against 5.23 in the real market over 2015 to 2025, and the 10-year 4.96 against 5.41. The index's daily correlation with a Treasury bond's return is -0.14 against a real -0.16 (IEF), and with an investment-grade bond's +0.20 against +0.27 (LQD). These are long-run rows R1 to R4. Under pt-v19 the curve is much quieter, 0.46 and 3.1 basis points a day, and bonds and stocks are uncorrelated. The library's tools/bonds/realism.py makes that comparison. Scenarios documents curve_shock.yml , which moves the whole curve 200 basis points in one day, and Agents and evaluation documents baselines.Balanced , a 60/40 book with a bond sleeve. The agent-facing book Engine.submit sends an order to the engine's own book under an agent's label, and the engine keeps each fill until take_fills collects it. Every share an order takes reaches the market once, on the next tick the market is open, and take_impacts reports the permanent impact it left. Portfolio.submit_limit and a tf.Limit in a World agent's orders both go through it. The table under Engine lists the seven methods. Seven ModelParams dials, listed on Model parameters, decide what the book holds. They are 0.0 on pt-v1 through pt-v19, where Engine.book_live is False: an order meets the maker's ladder exactly as Engine.book shows it, two agents trading one name in one step fill at the same prices because neither order removes a level, and a limit order's unfilled part waits outside the book, with mode range , and fills in full at its limit on the first tick whose print reaches it. pt-v20, the default, sets all seven, so Engine.book_live is True there and an order executes in the engine's own book. A waiting limit order then has mode queue . book_depth_coefficient , book_depth_exponent and book_depth_reach put latent depth beside the maker's ladder, where a larger order pays more per share, by the square-root law unless the exponent sets another power. book_shared makes an order consume what it takes. The maker's ladder is whole again when the maker quotes at the next tick, and the latent depth refills with a half-life of book_refill_half_life ticks. book_resting queues a limit order's unfilled part in the book, behind the depth already at its price, where the market's own flow or another agent can fill it. The flow fills it only at a price inside the maker's quote for that tick, so an order resting past the quote waits until the price comes to it. An order that the maker's next quote crosses trades at the maker's price, and its fill is recorded as liquidity="taker" . fill_impact_coefficient gives each agent's fills a linear permanent impact and attributes it to that agent. None of the seven changes a market nobody trades, and an engine no agent has sent an order to hashes and snapshots as it did under 0.8.1. The book holds equities only, so a rate index trades through Portfolio.execute on a book of its own. e.open_market() r = e.submit("alice", u[0].ticker, 1_000) r["filled"], r["average_price"] w = e.submit("alice", u[0].ticker, 500, limit_price=0.98 * r["reference"]) e.open_orders("alice") # mode "queue" on pt-v20 e.cancel(w["order_id"], agent="alice") e.take_fills("alice") # the market order's fill Portfolio Cash, positions and P&L for one trader. evaluate , TradingEnv , tca.analyse and World keep one per agent, and a loop of your own can use one the same way. execute fills an order at the prices the book gives and adds the trade to pending_flow , which the loop hands to the next session as run_session(fills=...) . Portfolio( cash: float = 1_000_000.0, *, max_leverage: float | None = None, cash_interest: bool = False, owner: str = "agent", margin_interest: bool = True, ) Signature Meaning execute(engine, ticker, quantity) -> dict Trade quantity shares, positive to buy, at the prices the book gives, and return the fill. A partial fill is reported as partial. Raises OrderError when nothing can fill or the trade would take leverage past max_leverage. submit_limit( engine, ticker, quantity, price, ) -> dict New in 0.8.5. Send a limit order to Engine.submit under owner and return the engine's report. What the book holds at price or better fills at once, and the rest waits until it fills or is canceled. max_leverage is checked as though the whole order filled at price, before anything is sent. A rate index raises ValidationError. cancel( engine, *, ticker=None, order_id=None, ) -> int New in 0.8.5. Cancel this portfolio's waiting orders, one by id, every one on a ticker, or all of them. Returns how many it canceled. open_orders(engine) -> list[dict] New in 0.8.5. This portfolio's waiting orders, in arrival order, as Engine.open_orders returns them. sync(engine) -> list[dict] New in 0.8.5. Apply every fill the engine holds for this portfolio, and return them. A waiting order that fills during a session reaches the portfolio this way, so a loop calls it after each session. It asks the engine nothing for a portfolio that has never sent an order to the book. pending_flow() -> dict[str, tuple[float, float]] Shares bought and sold per ticker since the last clear_flow. Pass it to the next run_session as fills=, or to a single tick as order_flow=, then call clear_flow, because flow left in place is sent again with the next step's. clear_flow() -> None Forget the accumulated flow once it has been passed on. accrue(engine) -> float New in 0.8.5. Book one trading day's interest, cash times the policy rate over 252, to cash and to interest, and return it. A positive balance earns it only with cash_interest on. A negative balance is charged it while margin_interest is on, the default, whatever cash_interest says, which is cheaper than any broker lends. evaluate, rank and World call it once a day, before the close. net_worth(engine) -> float pnl(engine) -> float Cash plus every position marked at the engine's prices, and that figure less starting_cash. leverage(engine) -> float gross_exposure(engine) -> float Gross exposure as a multiple of net worth, and the absolute market value of every position, longs and shorts alike. market_value(engine) -> float unrealised(engine) -> float realised() -> float The market value of the positions, their unrealised P&L, and the P&L realised on the parts already closed. marks(engine) -> dict[str, float] The current price per ticker, from the engine. fills_table(tickers) The fill log as an Arrow stream keyed by instrument_id, so it joins to bars and truth. stamp(day, step, tick) -> None Tag the fills that follow with a day, a step counted across the run and a tick within the day. owner: str New in 0.8.5. The label this portfolio's orders carry in the engine's book, "agent" by default. Portfolios that share one engine need different owners, and World gives each agent's portfolio that agent's label. cash_interest: bool interest: float New in 0.8.5. Whether accrue pays interest on idle cash, off by default, and the interest credited so far, net of any charged. margin_interest, on by default, is whether a negative balance is charged. cash, starting_cash: float positions: dict[str, Position] fills: list[dict] Cash now and at the start, each position by ticker, and every fill in the order it happened. When Engine.book_live is True, execute sends the order to Engine.submit under owner instead of pricing it off a copy of the book. The engine then applies the flow itself and pending_flow holds nothing for that trade, so a loop that passes fills=p.pending_flow() still counts it once. A rate index is always priced off its own book. import tradefloor as tf u = tf.Universe.random(20, seed=7) e = tf.Engine(seed=42, universe=u) p = tf.Portfolio(cash_interest=True) e.open_market() minute = 9 * 60 + 30 for step in range(6): # six 65-tick steps make one day if step == 0: p.execute(e, u[0].ticker, 500) e.run_session(minute // 60, minute % 60, 3, 65, fills=p.pending_flow()) p.clear_flow() minute += 65 p.accrue(e) # a day's interest, before the close e.close_market() p.net_worth(e) Exceptions Type Base Raised when ValidationError ValueError An input is refused, for example an unknown preset or parameter name, a universe that does not match a checkpoint's fingerprint, a rate passed as a percentage, or a count below one. OrderError ValueError The book cannot accept an order, for example one with a quantity of zero or less, an unknown side, or a cancel for an order id that is not resting in the book. Engine.submit raises it for an order the engine's book refuses, such as one on a rate index or under a label the book keeps for itself: mm, depth, flow or range. EngineBatch Runs one market under many seeds in lockstep, in one process. Build it like Engine, with seeds in place of seed. The run and read methods match Engine's and return one result per seed. It runs equities only and refuses a roster that holds a rate index. Run one on Engine, or through run_many, one seed per engine. seeds: list[int] tickers: list[str] draws_consumed: list[int] shape: tuple[int, int] model: ModelParams model_fingerprint: str EngineBatch(*, seeds, universe, ...) __len__() -> int open_market() -> None tick(...) / run_session(...) prices() -> bytes column(field) -> bytes News One news item that you inject into the market. News( *, ticker: str | None = None, sector: str | None = None, price_impact: float = 0.0, ) -> None NewsImpact The effect of one news item on prices. NewsImpact( *, ticker: str | None = None, sector: str | None = None, sectors: Sequence[str] | None = None, remaining_impact: float = 0.0, reversal_phase: bool = False, ) -> None TickResult What one tick did, as tick() returns it. market_status: str draws_consumed: int active: int run_many tf.run_many( seeds: Iterable[int], *, universe: Sequence[Instrument], macro: Macro | None = None, days: int = 1, ticks: int = 390, start: tuple[int, int, int] = (9, 30, 3), workers: int | None = None, collect: str = "prices", scenario: Any = None, model: str | ModelParams | None = None, ) -> list[Any] Runs one simulation per seed in a thread pool and returns the results in input order. collect chooses what comes back: "prices" (raw f64 bytes), "attribution" (the factor columns) or "summary" (prices, draw count, tickers and fingerprints). Every seed runs the same model. A single seed always runs in-process. Examples Forking and pinning e = tf.Engine(seed=42, universe=u) e.run_days(60) a, b = e.fork(2) # two engines, one past b.pin_macro(vix=45.0) a.run_days(20) b.run_days(20) # same noise, higher fear Reading a run back e.run_days(30, record=True) bars = e.bars(grain="day") # every recorded day one = e.truth(day=12) # that day alone # day=None means every recorded day; # an unrecorded day raises The book book = e.book(u[0].ticker) est = book.sweep_cost("buy", 50_000) # cost without executing; None on an empty side res = book.submit("buy", 50_000, taker="me") [f.price for f in res.fills] Universes u = tf.Universe.random(40, seed=7) u2 = tf.Universe.from_json(u.to_json()) u.fingerprint == u2.fingerprint # True # order is contractual: a reordered roster # is a different market from the same seed See also How prices are made explains how truth() breaks each price move into parts. Model parameters lists the coefficient set. Forks and counterfactuals covers saving, forking and resuming a market. ====================================================================== # Agents and evaluation https://docs.tradefloor.dev/evaluate.html Reference for tradefloor agents: what act returns, what an agent can see, and evaluate, rank, tca.analyse, Scorecard, Ranking, AgentRecord and RunManifest. ====================================================================== API REFERENCE/ AGENTS AND EVALUATION Agents and evaluation Reference for the agent protocol, tf.evaluate, tf.rank, the baselines, tf.tca.analyse and the records they return: Scorecard, Ranking, AgentRecord, Execution and RunManifest. Worked examples are in the guides: Write an agent builds an agent and sizes its orders, Compare strategies ranks agents across seeds and prices their trading, and Record and replay publishes a run a reader can check. Writing an agent An agent is any Python object with an act(obs) method. At each decision step the harness passes it an observation and reads back orders. By default a run has 6 decision steps a day, 65 simulated minutes apart, and each agent starts with 1,000,000 in cash and may hold positions worth up to twice its net worth. obs.step counts steps across the whole run, and obs.step_of_day, obs.is_first_step_of_day and obs.is_last_step_of_day count within the day. An agent may also declare privileged = True and keep an explain(day) method, both described below. A first agent builds and scores one on a seeded market. What act returns A dict of ticker to order, where each value is one of these: Value Order 1000 or -500 A market order for that many shares. It fills against the book at once, and what the book cannot fill is listed in partial_fills. tf.Limit(500, 101.25) Buy up to 500 shares at 101.25 or better, or sell with a negative quantity. What does not fill waits in the book, behind the shares already at that price, until it fills, you cancel it, or a new limit on the ticker replaces it. tf.Cancel() Cancel every order of yours still waiting on that ticker. Return None or {} to trade nothing. Values are shares, never weights, and any finite number of shares is accepted, fractions included. To hold a target weight, send the target holding less the shares already held and the shares in orders still waiting on the ticker, with the target rounded to whole shares if you want them: target = int(weight * obs.portfolio.net_worth() / obs.price(ticker)) order = target - obs.position(ticker) - waiting The target alone, such as 0.2 * net_worth / price, is the holding to reach, and sent as an order while a position exists it buys the position again. Sizing to a target weight works through an example with a waiting order. There are no stop or bracket orders, so check a stop yourself at each step. A refused entry, such as a string or NaN, gets a line in the scorecard's errors and the rest of the dict still trades, so read errors before the P&L. What the agent sees The observation holds the roster and last prices, each name's order book and average volume, the agent's own positions, cash and waiting orders, daily bars and the published economic figures in obs.history, and a view of the market in obs.engine. What obs.engine and obs.portfolio serve depends on the agent's access: ordinary by default, privileged with privileged = True on the agent, or trusted with trusted_agents=True on the call. The levels are the same under evaluate, rank, World and tca.analyse, and The agent's view sets out what each one serves. Ordinary and privileged agents get the read-only MarketView and PortfolioView, which raise tf.SandboxError for anything they do not serve, recorded in errors. obs.hidden is None unless the agent is privileged, when it is a read-only HiddenState that also carries the size of each day's news. obs.engine.bars() refuses under evaluate, which records no days, so daily bars come from obs.history. An LLM framework run through an adapter sees only the payload the adapter builds, the same at every level: What the model sees. Warm-up history history_days=N on evaluate, rank, World or tca.analyse runs the market for N days before day 0 with nobody trading, and obs.history holds those days at the first decision, labeled -N to -1. Each scored day is added after its last step. N is at most 2,520, ten years of 252 days, and no scenario applies during the warm-up. Warm-up history runs a breakout rule on it. Against buy-and-hold, across seeds evaluate scores one market, which is one sample. Compare each agent with buy-and-hold on the same market with versus_buy_and_hold, and run rank over many seeds before calling a winner. Compare strategies does both and reads the paired sign test. To put several agents in one market, use a World. Explaining moves An agent may also have explain(day). The harness calls it after each close and expects the name of the factor that moved prices most that day. On pt-v20 random_noise wins almost every day, so the scorecard reports explanation_baseline, what naming the most common winner every day scored, and explanation_edge, the accuracy less that baseline, which is the figure to quote. Execution cost tf.tca.analyse prices every fill against the same seed run without the agent's orders. Execution cost in Compare strategies prices a round trip and says what the figure measures, and tca.analyse below has the arguments. evaluate Score each agent on its own copy of the same market. tf.evaluate( agents: dict[str, Agent | StrategySpec], *, seed: int, universe: Sequence[Instrument], macro: Macro | None = None, days: int = 5, steps_per_day: int = 6, ticks_per_step: int = 65, cash: float = 1000000.0, max_leverage: float | None = 2.0, start: tuple[int, int, int] = (9, 30, 3), scenario: Scenario | None = None, model: str | ModelParams | None = None, cash_interest: bool = False, trusted_agents: bool = False, history_days: int = 0, margin_interest: bool = True, ) -> dict[str, Scorecard] Argument Default Meaning agents required Name to agent. An agent is any object with act(obs), or a StrategySpec, which is built fresh on every call. seed required Any integer from 0 to 2**64 - 1. Fixes the market every agent meets. universe required The roster. Its order is part of the market. macro None The economy on day 0. None uses Macro()'s defaults. days 5 Trading days to score. steps_per_day 6 Decisions per day. The agent acts once per step. ticks_per_step 65 Minutes between decisions. cash 1000000.0 Starting cash per agent. max_leverage 2.0 Cap on gross exposure as a multiple of net worth. None removes it. start (9, 30, 3) Hour, minute and day of the week the first session opens on. scenario None A Scenario to drive the economy. None lets it run on its own. model None A preset name or a ModelParams. None is the default preset, pt-v20. cash_interest False True pays uninvested cash the policy rate, a day's worth before each close. trusted_agents False True hands agents the live engine and portfolio in place of the read-only views. Every scorecard then says trusted. history_days 0 Days to run before day 0 with nobody trading, readable through obs.history. At most 2,520. The scored days continue the warmed market. margin_interest True A negative cash balance pays the policy rate before each close. False borrows for free, as 0.8.1 and earlier did, and the scorecard says free-borrowing. Returns a dict of Scorecard, keyed by the names you passed. evaluate builds one engine and forks it for each agent and for an untraded baseline, so every agent meets the same market and none uses liquidity another expected. Agent objects are not copied, so state an agent keeps persists across its own steps. An agent's fills in a step reach its market once, on the first tick of the next step. A bad entry in an agent's dict is refused with a line in errors and the rest trades. A return other than a dict, or an exception from act, trades nothing that step. evaluate warns, without changing anything it runs, when an agent failed on every step, when every order in a step was a fraction of a share (weights sent as shares), and when a spec's top_k is more than half the roster. reference_agents The five baseline agents. tf.reference_agents(*, seed: int = 0) -> dict[str, Agent] seed seeds the random agent only. The keys are buy_and_hold, random, momentum, mean_reversion and oracle. The Oracle declares privileged = True and reads the model's fair value, holding the same gross exposure as the others. On pt-v20 it is a reference agent and not a ceiling, so compare agents with buy-and-hold there. baselines.Balanced A fixed-weight portfolio of the equities and the simulated rate indices, 60/40 by default, with a drift band. It is one of the baselines reference_agents leaves out. tf.baselines.Balanced( *, equity: float = 0.6, bonds: dict[str, float] | None = None, band: float | None = 0.05, max_participation: float = 0.05, ) Argument Default Meaning equity 0.6 Share of net worth held across every equity in the roster, equally weighted. bonds None Rate index to share of net worth. None is UST2Y 0.10, UST10Y 0.20 and IGCORP 0.10. band 0.05 How far a weight may drift before the agent trades every holding back to target, checked at each day's first step. None never rebalances. max_participation 0.05 The largest trade in one name as a share of its average daily volume. It needs a roster with the rate indices, from Universe.random(..., bonds=True) or Universe.with_bonds(). Without them it raises ValueError at its first step, which evaluate records in errors. u = tf.Universe.random(20, seed=101, bonds=True) scores = tf.evaluate( {"sixty_forty": tf.baselines.Balanced()}, seed=3, universe=u, days=120, scenario=tf.Scenario.load("curve_shock"), cash_interest=True) capture_ratio tf.versus_buy_and_hold(scores, *, reference: str = "buy_and_hold") -> dict[str, float] tf.capture_ratio(scores, *, oracle: str = "oracle") -> dict[str, float] tf.capture_withheld(scores, *, oracle: str = "oracle") -> str | None tf.oracle_is_ceiling(model=None) -> bool versus_buy_and_hold gives each agent's P&L less buy-and-hold's in the same market, in currency. It is the comparison to quote on pt-v20. It leaves out tampered agents and warns, returning {}, when buy-and-hold did not run. capture_ratio gives each agent's P&L as a fraction of the Oracle's: 1.0 is the Oracle's P&L. On pt-v20 it returns {} and warns, because a shock there moves fair value for good and the Oracle's P&L follows the market's month. capture_withheld returns that reason, or None where a ratio is reported. oracle_is_ceiling answers for a preset name, a ModelParams or the default. leaderboard tf.leaderboard(scores, by: str = "pnl") -> list[Scorecard] Sorts scorecards by one field, best first, with tampered cards last. rank Run the same agents on many seeds and compare them seed by seed. tf.rank( make_agents: Callable[[], dict[str, Agent | StrategySpec]], *, seeds: Iterable[int], universe: Sequence[Instrument], macro: Macro | None = None, days: int = 5, steps_per_day: int = 6, ticks_per_step: int = 65, cash: float = 1000000.0, max_leverage: float | None = 2.0, start: tuple[int, int, int] = (9, 30, 3), scenario: Scenario | None = None, oracle: str = "oracle", workers: int = 1, model: str | ModelParams | None = None, trusted_agents: bool = False, benchmark: str = "buy_and_hold", history_days: int = 0, margin_interest: bool = True, ) -> Ranking The arguments shared with evaluate mean the same, and the others are these: Argument Default Meaning make_agents required Called once per seed, and must return new agents each time, because an agent carried across seeds brings its state with it. seeds required The seeds to run. Each is one market. oracle "oracle" The entrant to measure capture against, where capture is reported. workers 1 Threads to spread the seeds over. 1 runs them one after another. Results are collected in seed order either way. benchmark "buy_and_hold" The entrant each agent's excess P&L is measured against. Ranking Member Meaning records One AgentRecord per agent, keyed by name. seeds The seeds that ran. benchmark, benchmark_note The entrant excess P&L is measured against, and how it was chosen. capture_withheld Why no capture ratio is reported, on pt-v20. None where one is. tampered Agents left out because they changed the market. oracle, reference_pnls, unmeasurable The capture reference, its P&L per seed, and the seeds where no capture could be formed. model_fingerprint, universe_fingerprint The model and roster every seed ran on. report(), table(), as_dict() The results as a printable table, as rows, and as plain data. separation(a, b) A paired sign test between two agents. report() names the benchmark by its label, as in vs buy_and_hold or vs flat. On pt-v20 it also says that the Oracle ran and has no row, with its P&L over the benchmark's. An agent whose adapter refused actions gets a REFUSED line with the count and the first refusal, an agent whose act() returned something other than a mapping gets an UNUSABLE line with the count and the first seed and step, and an agent that raised gets a RAISED line. separation(a, b) returns a dict with a, b, wins_a, wins_b, ties, paired_seeds, decisive and p_value. Both agents meet the same market on each seed, so each pair differs only in the agent. decisive is true only when one agent won on every seed, and it counts wins without testing significance. AgentRecord One agent's results across every seed. Member Units Meaning name, seeds The name it ran under, and its seeds. pnls currency P&L per seed. median_pnl currency The median of pnls. excess_pnls, mean_excess_pnl currency P&L less the benchmark's per seed, and their mean. The number to quote on pt-v20. seeds_ahead seeds Seeds on which the agent earned more than the benchmark. wins, win_rate seeds, fraction Seeds on which it had the highest P&L of the ranked agents, and that as a share. pooled_capture, median_capture, capture_range, captures ratio Capture of the Oracle's P&L, pooled, per seed and its spread. None or empty on pt-v20. errors, seeds_with_errors, first_error How often the agent's code raised on each seed, on how many seeds, and the first line. rejected, max_leverage orders, multiple Refused orders and peak leverage per seed. refused actions Per seed, the actions an LLM adapter turned down on its own, such as a ticker the roster does not have, while the rest of the decision traded. They are part of rejected and are not raises. first_refusal The first of those refusals, prefixed with its seed, or None. unusable, first_unusable steps Per seed, the steps whose act() returned something other than a mapping, such as a list of pairs, and the first of them. They are not counted in errors. failed_every_step, seeds_failed Per seed, whether every step raised or was refused and nothing traded, and on how many seeds. Such a seed has no score: it reads None in excess_pnls, does not count in seeds_ahead, and cannot win. trusted, uses_hidden_state How this agent saw the market. as_dict() The record as plain data. tca.analyse Run the same seed with and without the agent's orders, so every fill is priced against the market where it never traded. tf.tca.analyse( agent, *, seed: int, universe: Sequence[Instrument], macro: Macro | None = None, days: int = 1, steps_per_day: int = 6, ticks_per_step: int = 65, cash: float = 1000000.0, max_leverage: float | None = 2.0, start: tuple[int, int, int] = (9, 30, 3), scenario: Scenario | None = None, model: str | ModelParams | None = None, trusted_agents: bool = False, history_days: int = 0, ) -> Execution The agent is any object with act(obs) returning share quantities, as on Writing an agent. tca.analyse takes market orders only and refuses a tf.Limit, because the part of a limit order that waits fills inside a session, where the untraded run has no price to compare it with. days is 1 by default, unlike evaluate. history_days runs the market that many days before day 0 with nobody trading, as in evaluate, and both runs start day 0 from the warmed market. No scenario applies during the warm-up. Execution cost explains what the comparison measures. Execution What the trading cost, against the run without the agent's orders. Member Units Meaning shortfall_bps() basis points Cost of trading against the untraded path. Positive is a cost. shortfall currency The same in currency. by_step(), by_ticker() The cost per decision step and per name. partial_fills() What was asked for against what filled. impact_bps(ticker) basis points One name's price move caused by the agent's own orders. fills Every fill the agent took. actual_final, baseline_final currency Net worth at the end, traded and untraded. actual_path, baseline_path currency Net worth per step, traded and untraded. moved, untouched_moved() Names whose price the agent moved, and names it never traded whose price still differs from the baseline. steps, tickers, portfolio, seed The run's size, roster, final portfolio and seed. history_days days The warm-up both runs had before day 0. model_fingerprint, universe_fingerprint The model and roster it ran on. as_dict() The result as plain data. Scorecard One agent's result on one seed, as evaluate returns it, with errors to read before the P&L. Attribute Units Meaning name, seed The name it ran under and the seed. pnl currency Final net worth less starting cash. return_pct percent P&L as a percentage of starting cash: 5.2 means 5.2%. final_net_worth currency Cash plus positions marked at the last close. equity_curve currency Net worth after each day's close. The last value is final_net_worth. max_drawdown_pct percent The largest fall from a running peak of equity_curve, which starts at the starting cash. Measured on closes. sharpe ratio Mean daily return over its standard deviation, times the square root of 252, from equity_curve and the starting cash. No risk-free rate is subtracted. None with fewer than two days, once net worth reached zero, or when the returns did not vary. Over five or ten days it is very noisy. volatility_pct percent Annualized volatility of the same daily returns: their standard deviation times the square root of 252. None with fewer than two days or once net worth reached zero. exposure_curve multiple Gross exposure as a multiple of net worth after each step's session, longs and shorts both counted. Infinite once net worth is at or below zero. time_in_market fraction Share of steps in exposure_curve that ended holding any position. avg_gross_exposure multiple Mean of exposure_curve over the steps where net worth was above zero. 1.0 is fully invested and 2.0 the default leverage limit. ruined Net worth was at or below zero at some close. impact_bps basis points The agent's own footprint, against the run where it did not trade. trades fills Fills taken. A market order, or the part of a limit order that fills on arrival, counts once however many price levels it took, and each later fill of a waiting limit order counts again. turnover currency Traded value. max_leverage multiple Highest gross exposure reached, as a multiple of net worth. rejected orders Orders refused, and actions an LLM adapter refused on its own. Each has a line in errors. leverage_refusals orders How many of rejected the leverage limit refused. partial_fills One line per market order the book could not fill in full. errors One line per step where the agent raised, returned something that is not an order dict, or sent a refused order. explanations What the agent's explain(day) answered, day by day. explanation_accuracy fraction Share of days on which explain named the factor that moved prices most. None without explain. explanation_baseline fraction What naming the most common winner every day scored on the same days. explanation_edge fraction Accuracy less baseline. On pt-v20 the baseline is 0.95 to 1.0, so this is the figure to compare. history_days days The warm-up before day 0. Left out of as_dict() when 0. margin_interest Whether borrowing paid the policy rate. trusted, uses_hidden_state, tampered The agent had the live engine, read hidden state, or changed the market outside its orders. A tampered card ranks nothing. strategy_fingerprint sha256 of the StrategySpec, or empty for a hand-written agent. model_fingerprint, universe_fingerprint The model and roster it ran on. as_dict() The scorecard as plain data. sharpe, volatility_pct, time_in_market and avg_gross_exposure are read-only properties. They and exposure_curve are left out of as_dict(), so the traded digest hashes what it did before they were added. The repr prints the four figures after the impact, then the counts and flags that are set. Over fewer than Scorecard.SHARPE_MIN_DAYS (20) scored days it prints sharpe=n/a (short run), because the standard error of an annualized Sharpe ratio is 3.5 or more there. The property still returns the figure: buy-and-hold in One market and the baselines reads sharpe=n/a (short run), vol=9.6%, in_market=100%, exposure=1.00x after ten days on a 1.05% gain, and its sharpe is +2.78. The counts and flags include errors=1, free-borrowing, history_days=20 and explanation=0.967 vs baseline 0.983 (edge -0.016). RunManifest A finished run as one document: the roster, the day-0 economy, the scenario, the order log, the strategy when it is a StrategySpec, the package version, the model's coefficients and the market digest a replay has to reproduce. Recording the market shows the round trip. tf.RunManifest.of( engine: Engine, *, seed: int, universe: Sequence[Instrument], macro: Macro | None = None, scenario: Scenario | None = None, strategy: StrategySpec | str | None = None, universe_source: Any = None, label: str = "", derived_from: Any = None, ledger: DayLedger | None = None, agent_access: dict[str, Any] | None = None, ) -> RunManifest Member Meaning of(engine, ...) Capture a finished run. The seed and roster are passed because an engine keeps neither. strategy takes a StrategySpec, carried in full, or a reference string for a hand-written agent, recorded in gaps. derived_from takes the Checkpoint the run branched from. to_json(), from_json(text) The manifest as JSON text, and back. reproduce() Replay and verify. Returns the rebuilt engine. verify_lineage(checkpoint) Check that derived_from names this checkpoint. describe() A readable summary. seed, label, universe, universe_source, macro, scenario The inputs as recorded. strategy, strategy_reference The carried spec, or the reference string of one that is not carried. order_log Every input the engine consumed. model, model_fingerprint The coefficients, and a preset's name or custom- plus a digest. fingerprints A digest per component, and inputs over the seed and all of them. result The market digest, the days run and the draws consumed. written_by The package version, platform, Python version, model and era digest of the build that wrote it. derived_from The checkpoint this run branched from, or None. gaps, complete What a reader needs from outside the manifest, and whether that list is empty. reproduce() raises ValidationError on a mismatch and names the component that disagreed. It first compares the era digest, a small fixed calculation, with the one in the manifest, and refuses to replay on a build that does different arithmetic. It checks the market and carries no score: result holds the market's digest, the number of days and draws_consumed, and nothing else. evaluate and rank write no manifest, so to let a reader check a score, publish the agent, the call and the seeds. ====================================================================== # Forks and counterfactuals https://docs.tradefloor.dev/api-counterfactual.html Fork a market, checkpoint it and run one agent in two worlds that differ in one input: Checkpoint, branch, World, agree and compare, with usage first. ====================================================================== API REFERENCE/ FORKS AND COUNTERFACTUALS Forks and counterfactuals Reference for World , agree , compare , their result types, and the tools for checkpoints and replay. A counterfactual runs one agent in two copies of a market that differ in one variable. Checkpoints and forks A fork copies a running market so you can run two futures from one past. Everything before the fork is identical in both copies, bit for bit, so a difference that appears afterwards comes from the input you changed. import tradefloor as tf universe = tf.Universe.random(40, seed=7) engine = tf.Engine(seed=42, universe=universe) engine.run_days(60) mark = tf.Checkpoint.of(engine, universe=universe, seed=42) text = mark.to_json() # outlives the process calm, hiked = mark.branch(2) # two engines, identical up to day 60 hiked.pin_macro(corporate_bond_yield=0.09) # they diverge only from here calm.run_days(20) hiked.run_days(20) Cost Survives the process tf.branch(engine, 2) under 1 ms no Checkpoint.resume() seconds, growing with the order log yes branch copies the engine state, every column and each random stream's position. Checkpoint stores the order log and rebuilds the engine by replaying it, so cite the checkpoint in a published result. A checkpoint records the roster's fingerprint and refuses to load against a roster that changed, because two rosters can share tickers and differ in fundamentals. Two futures from one past runs a fork with its output. Counterfactuals A World holds the market, one agent, its portfolio and the macro path. Forking a world copies all of them, so the same agent meets two markets that differ in one input: class Mine: def act(self, obs): return {"AAA": 300} if obs.step == 0 else None u = tf.Universe.random(24, seed=7) world = tf.World(seed=7, universe=u, agent=Mine()) world.run(50) # a shared past control, stress = world.fork("control", "stress") stress.apply(tf.Scenario.load("liquidity_crisis")) started = tf.agree(control, stress) # identical, bit for bit control.run(80) stress.run(80) result = tf.compare(control, stress, agreement=started) print(result.render()) agree checks that two arms match in engine state, order books, portfolio, agent state and stream positions. compare finds the first step where they differ, separately for the macro path, the agent's decisions, its orders, the prices and the portfolios, and summarizes both arms. world.manifest() records the run like any other, and an external agent runs here the same way, through an adapter. A counterfactual with an agent runs this example with its output and reads the comparison. The answer describes this simulated market under the one change you made. A claim about a real market needs evidence from a real market. World A market, the agent or agents trading in it, and the macro path they run under. It is available at the top level as tradefloor.World . World( *, seed: int, universe: Sequence[Instrument], agent: Any = None, agents: dict[str, Any] | None = None, pins: dict[str, Any] | None = None, macro: Macro | None = None, cash: float = 1_000_000.0, max_leverage: float | None = 2.0, steps_per_day: int = 6, ticks_per_step: int = 65, start: tuple[int, int, int] = (9, 30, 3), model: str | ModelParams | None = None, label: str = "", on_refusal: str = "raise", ) run(days=1) Advance the world by the given number of days. Each day runs steps_per_day decision steps, the points where the agent decides, and each step runs ticks_per_step ticks. Returns self. act(obs) orders An agent's act(obs) returns a dict of ticker to shares, positive to buy. A number is a market order, which the agent's portfolio fills through Portfolio.execute. From 0.8.5 a value can instead be tf.Limit(quantity, price), which takes what the book holds at price or better and leaves the rest waiting, or tf.Cancel(), which cancels the agent's waiting orders on that ticker. A new Limit on a ticker replaces the order the agent has waiting there. A Limit on a rate index is refused and recorded in the row's refused list. From 0.8.5 evaluate and rank take them too, and tca.analyse refuses a Limit. book_fills New in 0.8.5. A waiting order can fill during a session, while no agent is being asked. The fill reaches the agent's portfolio after that session through Portfolio.sync, and the step's trace row lists it under book_fills, per agent in a cohort. A row has book_fills only when a waiting order filled. agents={label: agent} Runs several agents in this one engine in place of agent, each in its own portfolio with its label as the portfolio's owner. Every agent starts with the same cash, in its own account: cash is one number, and cash={'a': 1e6, 'b': 5e6} is refused. max_leverage is one setting too, applied to each agent's book on its own. They are asked in label order, and their fills are summed and reach the market once, as the session's fills. With Engine.book_live False, as on pt-v1 through pt-v19, agents in one step meet the same book and take no levels from each other. On pt-v20, the default, book_shared is on, so label order is also arrival order: a later agent meets the book an earlier one left and can fill against its waiting limit order. Two identical buyers of 10% of a day's volume paid 19.5 to 30.7 bp apart on seeds 1 to 10, the later label paying more, so rotate the labels across runs when you compare agents in one market. on_refusal="raise" "raise" ends the run when an agent cannot produce a decision. "skip" records the refusal, trades nothing that step and carries on. tf.externalities(world, days=1) Measures what each agent in a cohort did to the others, without advancing world. It forks the cohort once per agent plus once, freezes one agent in each arm, runs every arm for days and returns an Externality. matrix[a][b] is the change in b's P&L when a stops trading. From 0.8.5, levels[a][b] is the change in b's execution cost against each step's opening mid when a stops trading, positive when a made b's fills dearer, and live records Engine.book_live. levels is zero off its diagonal wherever the book is not shared. A frozen agent's waiting orders are canceled in its arm. fork(*labels) Returns one World per label. Each copy, called an arm, matches this world at the fork in its engine, agent, portfolio and random generator position. Each arm records the step it was forked at. apply(scenario) Apply a Scenario document to this arm, shifted onto this world's own day numbering. Returns self. intervene(**fields) Change macro fields in this arm on the current day. The change is recorded next to the day-zero pins, so the world's scenario can still be rebuilt. checkpoint(label="") A Checkpoint of the run so far, which you can save and resume later. manifest(*, strategy=None, label="") A RunManifest, the record of a run's inputs, so others can cite and replay the counterfactual. replay() Re-execute the order log into a fresh Engine and return it. summary(*, since=None) / net_worth A summary of results, optionally from a given step onward, and the portfolio's current value. day / step / digest The next day to run, the next decision step counted across the whole run, and a hash of the run state. scenario / firings / order_log The reconstructed Scenario for this world, the interventions that fired, and every input that reached the engine. agree and compare tf.agree(a: World, b: World) -> Agreement tf.compare( control: World, treatment: World, *, agreement: Agreement | None = None, ) -> Comparison agree Checks that two arms are identical in engine state, order books, portfolio, agent state and random generator position. Call it at the fork, before the arms run on. compare Puts two forked worlds side by side and finds, for each series, the first step at which the two arms differed. The series are the macro values, the decision, the orders, the prices and the portfolios. Arms with different steps_per_day are refused. Pass agreement= so the published comparison states that the arms started identical. Agreement identical: bool, differences: list[str], as_dict(), render(width=22). Comparison Per-series Divergences and both arms' summaries. as_dict(), render(width=24). Divergence Where one series came apart: the step, the day and the values on each side. as_dict(), render(). Checkpointing and replay Checkpoint.of(engine, *, universe: Sequence[Instrument], seed: int, macro: Macro | None = None, label: str = "") -> Checkpoint tf.branch(engine, count: int = 2, *, universe: Sequence[Instrument] | None = None, seed: int | None = None, macro: Macro | None = None) -> list[Engine] tf.replay(log, *, seed: int, universe: Sequence[Instrument], macro: Macro | None = None, model: str | ModelParams | None = None, until: int | None = None) -> Engine Checkpoint.of Captures an engine's history. universe and seed are required because an engine is built from them and keeps neither. Checkpoint.resume / branch(count=2) Re-execute into one engine, or into count independent engines, at the saved point. Each build has an era digest that identifies its arithmetic. A checkpoint written by a build with a different era digest is refused, and the error names both builds. The digest changed in 0.8.0, so a checkpoint written by 0.7.x is refused even when it names its preset. Resume it under the version that wrote it. Checkpoint.to_json / from_json The serialized form, which you can save and load in a later process. Checkpoint.fingerprint A sha256 hash of the checkpoint, serialized in its standard form. RunManifest.of(..., derived_from=checkpoint) records it as the checkpoint the run came from. tf.branch Copies a running engine in constant time through Engine.fork. The copy includes every column, the random generator position, the news the engine generated that day, the tape and the order log. tf.replay Re-executes a recorded order log. seed and universe are not in the log, so you must pass them. See also Scenarios builds what World.apply takes, and RunManifest documents what a manifest carries. ====================================================================== # Scenarios https://docs.tradefloor.dev/api-scenario.html Apply a shipped scenario or write your own: the seven scenarios, the fifteen targets, the command line, then the Scenario and Intervention reference. ====================================================================== API REFERENCE/ SCENARIOS Scenarios Reference for Scenario , Intervention , Firing , the target registry and run_scenario . Using a scenario A scenario is a named set of changes to a market, with the assumptions behind them stated. Apply one to see how an agent behaves under a change you control, against the same market without it. import tradefloor as tf scenario = tf.Scenario.load("rate_shock") print(scenario.describe()) # shocks, then the author's assumptions u = tf.Universe.random(20, seed=101) scores = tf.evaluate(tf.reference_agents(seed=3), seed=7, universe=u, days=60, scenario=scenario) Seven scenarios ship with the package: curve_shock, geopolitical_conflict, liquidity_crisis, oil_price_spike, policy_regime_shift, rate_shock and recession. Each file keeps what happened (the shocks) apart from what the author assumes it caused (the transmission), and describe() prints the two under separate headings. tradefloor does not decide what a war or an oil shock does to a market. The file states it, and the engine runs the market with and without it. recession and liquidity_crisis were calibrated against 2008 and March 2020. On pt-v20, paired against the same seeds without the scenario, the recession takes the index down 44.7% at 120 sessions, against 45% from Lehman to March 2009. A scenario applies to an evaluation, a run loop, or one arm of a fork. Its days count from the day it is first applied, so on a branch at: 50 means fifty days after the fork. Writing one A scenario written in Python and the same one written in YAML give the same document and the same fingerprint, a sha256 of the resolved scenario that a RunManifest records. crisis = (tf.Scenario(name="my_liquidity_crisis") .shock("market.liquidity", operation="multiply", value=0.40, at=20, duration=25) .assume("macro.corporate_yield", operation="add", value=0.005, at=20, duration=25)) path = tf.Scenario().hold(vix=15.0).ramp("vix", start=48.0, end=22.0, over=45, begin=60) Operations are set, add and multiply; shapes are impulse, permanent, hold and ramp. An impulse writes the value once and lets the model carry it from there. A pin (hold, ramp, step) gives a whole path for one macro field. The targets A scenario can move fifteen targets: macro.corporate_yield, macro.policy_rate, macro.treasury_2y, macro.treasury_10y, macro.vix, macro.inflation, macro.growth, macro.unemployment, macro.cycle, macro.qe_pe_boost, macro.fear_greed, market.liquidity, market.earnings, commodity.oil and policy.tariff_rate. A name that looks like a target and is not one, such as market.volatility or execution.market_impact, is refused with an error that names what to use instead. tradefloor scenario targets prints each target with a note. The command line The scenario commands read scenario files and never run a market. tradefloor scenario list tradefloor scenario validate my.yml tradefloor scenario show oil_price_spike tradefloor scenario diff rate_shock recession tradefloor scenario targets Scenario A macro path, meaning day-by-day values for economy-wide series such as VIX, plus a set of interventions. The scenario is applied one day at a time. It is available at the top level as tradefloor.Scenario . Scenario( label: str = "", *, name: str | None = None, description: str = "", interventions: Sequence[Intervention] = (), shocks: Sequence[Intervention] = (), transmission: Sequence[Intervention] = (), vix_sets_variance: bool = False, ) hold(**fields) Pin fields to a constant from day zero. Returns self. scenario.FIELDS lists the fifteen fields you can pin, and any other name raises ValidationError. Fourteen are macro series, including treasury_yield_2y and treasury_yield_10y from 0.8.5, the yields the simulated rate indices read. The fifteenth, epicentre, names the sector that carries the next crisis episode, or "none" for a crisis with no such sector, as in hold(vix=65.0, epicentre="financial_services"). A misspelled sector raises an error in the call that names it. A pinned epicentre uses no random draw. It has no effect on a preset with crisis_epicentre_extra at 0.0, which is every preset before pt-v19. ramp(field, *, start, end, over, begin=0) Move a field in a straight line from start to end over the given days. Before begin the field stays at start, and afterward it stays at end. over < 1 or begin < 0 raises. step(field, *, before, after, at) Jump a field from before to after on day at. Several pins on one field run as back-to-back segments, and each must start later than the one before or the call raises. shock(target, *, operation="multiply", value=None, at=0, duration=None, shape=None) Add an intervention for an event from outside the market. Returns self. assume(...) Works like shock and takes the same arguments, but files the intervention as assumed transmission. describe() prints the two under separate headings. intervene(intervention) Add one Intervention you have already built. The order you add them in is kept, and it decides which of two same-day interventions goes first. apply(engine, day) Apply pins, then shocks, then transmission to one day of one run. Returns the list of Firings. Days count from when the scenario is first applied. at(day) / table(days) The pinned values for one day, or for the whole path, so you can check them before running. load(name) Classmethod. Loads a shipped scenario by name: curve_shock, geopolitical_conflict, liquidity_crisis, oil_price_spike, policy_regime_shift, rate_shock, recession. from_yaml(source) / from_json(text) Classmethods. Parse a scenario document. Validation errors name the target and the operation. rate_shock(*, start=0.025, end=0.05, over=30, begin=0, credit_spread=0.02) Constructor. Moves the policy rate and the corporate yield together, keeping them credit_spread apart. Raises if called on an instance. vix_shock(*, calm=15.0, peak=45.0, at=10, over=20) Constructor. A VIX spike that steps up from calm to peak on day at, then ramps back to calm over the given days. document() / fingerprint The resolved document in its standard form, and the sha256 hash of it. The YAML and Python forms of one experiment have the same fingerprint. vix_sets_variance appears in the document only when it is on, so a scenario written before 0.8.0 keeps its fingerprint. describe() The scenario as text, shocks above assumptions. log / firing_table() The last run's record of each firing, with the values it saw. without_interventions() / copy() The same macro path with interventions removed, and an independent copy for driving two runs at once. name / description / source The scenario's name, its description, and the file name it was read from, if any. vix_sets_variance Whether a VIX that this scenario forces also sets the market's volatility. It is read-only, and False unless the constructor, a YAML scenario: block or a to_json document turned it on. See below. A forced VIX that sets volatility A scenario forces the VIX when it sets it with a pin or an intervention. By default a forced VIX reaches volatility the same way the model's own VIX does. Each close moves the variance of the market factor, the part of every price move that all companies share, one step toward the level that VIX implies. The fast component closes about 2% of the gap each trading day, a half-life of about 33 trading days. So a VIX that jumps from 15 to 80 in three weeks shows up in prices weeks later. Scenario(vix_sets_variance=True) changes that. On every trading day that the scenario forces the VIX, with a pin or an intervention on macro.vix , the close sets both variance components to the level the variance law (the model's rule for how variance moves) reverts to at that VIX, clamped as usual. The next day trades at that level. While the VIX is forced, the day's own market shock does not feed into the factor's variance. Volatility for each company and each sector, and jumps, still cluster on their own shocks. When the scenario stops forcing the VIX, the variance law carries on freely from that level with no jump. When the real 2020 VIX is replayed on pt-v19, the model's worst month comes 2 trading days after the real one with the switch on, and 20 trading days after it with the switch off. In YAML the switch is vix_sets_variance: true in the scenario: block. In a to_json document it is "vix_sets_variance": true beside "path" . It must be a boolean, so a quoted "true" is refused. A scenario that turns the switch on but never forces the VIX raises ScenarioValidationError when it is applied, because the switch would do nothing. With the switch off, every earlier scenario runs, serializes and fingerprints as before. At the engine level the switch is Engine.pin_macro(vix=..., vix_sets_variance=True) , which marks the current day's close. Engine.vix_sets_variance_pending reports the mark until that close clears it. Intervention One change to one target at one time. An Intervention is immutable and is checked when you build it, and it is available at the top level as tradefloor.Intervention . Intervention( target: str, *, operation: str = "multiply", # "set" | "add" | "multiply" value: Any = None, at: int = 0, # days from scenario start duration: int | None = None, shape: str | None = None, # "impulse" | "permanent" | "hold" | "ramp" role: str = "shock", # "shock" | "transmission" ) as_dict() / from_dict() The serialized form, with every default filled in, so the fingerprint is the same whether or not you typed the defaults. active_on(day) / last_day Whether the intervention acts on a day, and the last day of its window. describe() One line, in the units the target declares. ScenarioValidationError Raised for a malformed scenario document or an unknown target. A subclass of ValidationError. Firing One application of one intervention, with its target, its day, the value read and the value written. as_dict() returns a dict and str() returns readable text. The target registry tradefloor.TARGETS maps the fifteen target names to Target objects. macro.treasury_2y, macro.treasury_10y and market.earnings joined in 0.8.5. tradefloor.UNSUPPORTED_TARGETS maps each refused name to what to use instead. The targets, above, lists the names, and tradefloor scenario targets prints each with a note. Target.read(engine) / write(engine, value) How the target reads and writes the engine. Target.check(operation, value) The checks run when an intervention is built. Rates must be fractions in [-0.05, 0.50], prices must be positive, and a cycle phase must be given by name. Target.show(value) The value in the target's own units, for messages and audit trails. run_scenario tf.run_scenario( scenario: Scenario, *, seed: int, universe: Sequence[Instrument], days: int, macro: Macro | None = None, ticks_per_day: int = 390, start: tuple[int, int, int] = (9, 30, 3), record: bool = False, model: str | ModelParams | None = None, ) -> Engine Runs a market under the scenario and returns the finished engine. The scenario is applied at the start of each day, so the path already applies on day zero. days < 1 raises ValidationError. tf.evaluate , tf.rank , tf.tca.analyse and tf.run_many take the same scenario= keyword. See also Forks and counterfactuals for World.apply. Agents and evaluation for the evaluation entry points that accept a scenario. ====================================================================== # LLM adapters and MCP https://docs.tradefloor.dev/api-integrations.html Run an agent from LangGraph, the OpenAI Agents SDK, PydanticAI, FinRobot or a plain function, record and replay it, and use the MCP server's tools. ====================================================================== API REFERENCE/ LLM ADAPTERS AND MCP LLM adapters and MCP Reference for the LLM adapters and the local MCP server. The guides set each one up and run it: LLM adapters, Local MCP server and Record and replay. A summary of each comes first, then the shared layer and the four framework adapters: signatures, input and output mapping, errors, replay, tracing and the framework version each was written against. Bringing an LLM agent An adapter runs an agent built in another framework inside a tradefloor market. The framework reads a JSON observation and answers with a decision, and tradefloor checks the decision, sends the orders to the book and scores the result. Each adapter is an ordinary agent with an act method, so it runs under tf.evaluate, tf.rank and World, and two frameworks run on the same seed meet the same market. LLM adapters sets each framework up and runs it with no API key, and Record and replay records a run and replays it without the model. Framework Module Install extra Plain Python function tradefloor.integrations.callable none OpenAI Agents SDK tradefloor.integrations.openai_agents openai-agents PydanticAI tradefloor.integrations.pydantic_ai pydantic-ai LangGraph tradefloor.integrations.langgraph langgraph FinRobot tradefloor.integrations.finrobot finrobot, Python 3.11 only An adapter asks its framework every every decision steps, 6 by default, which at the default 6 steps a day is one decision a simulated day. One decision can take several calls to the model provider, through tool calls, framework turns and retries, as Decisions and model calls sets out. The decision is the same for every adapter: a list of actions, each with a symbol, a side (BUY, SELL, HOLD or CANCEL), a positive quantity, and an optional limit_price that makes it a limit order. A bad action, such as an unknown symbol or a negative quantity, is refused on its own and recorded in the scorecard's errors, and the rest trade. An order over the participation cap, 2% of the name's average daily volume by default, is cut to the cap. The observation payload never carries fair value, the attribution of a move or the macro path the run has not reached. A seed replays the market exactly, and a Transcript replays the model. Each entry is keyed by the payload the model was shown, so a changed roster, seed or cadence stops a replay with ReplayMiss, and a replay under different instructions is refused before the market opens. To publish a result from an external agent, record both sides: adapter.provenance() for the framework, model and settings, and a RunManifest for the market. tf.fingerprint.fingerprint(agent) runs an agent on a fixed battery of seven 120-day markets, one per shipped scenario, and hashes what it ordered, so two versions of an agent can be checked for whether they behave the same. tf.battery() returns that battery and BATTERY_VERSION names it. To show a score was not tuned to its seeds, publish tf.commit(seeds, salt) before the run, then the seeds and the salt, which tf.reveal checks. tf.sealed_battery(seeds, salt) builds the battery on them. The MCP server tradefloor-mcp offers the simulator as tools a model calls over the Model Context Protocol, on your machine over stdio. Local MCP server installs it, registers it with a client and makes a first call. Over MCP a model studies the market from outside. No tool places an order, and a tool that runs a market gives the same result for the same arguments. A model that trades inside the market is an agent, run through an adapter, and the hosted app serves its own MCP tools, which do place orders. Tool What it does describe_simulator what the simulator is, what it is measured to reproduce and what it cannot do check_envelope whether a question is inside the validated scope, before anything runs validate_strategy checks a strategy spec and returns its fingerprint build_universe builds a roster: generated, concentrated in some sectors, or written by hand build_scenario composes a scenario and shows what it resolves to list_scenarios the shipped scenarios, the constructors and the targets evaluate_strategies runs strategies and the five baselines on one market rank_strategies runs them across seeds, with a paired sign test between each pair run_stress_scenario what a scenario does to each strategy, against the same market without it explain_price_move a price move split into the eleven factors that sum to it explain the random draws behind one day for one company start_job, check_job slow work in the background, up to 252 days A strategy is a JSON StrategySpec, never code, and every result carries a caveats list and a provenance block with the version, preset, seed and roster, so a model summarizing it sees the limits too. A direct call runs at most 60 days and a background job 252, the horizon the model is validated over. No tool takes a preset: every run uses pt-v20 and names it. Limits lists the rest of the caps. The shared layer tradefloor.integrations.common holds the parts every adapter needs that belong to no single framework. These are the observation allowlist, the decision schema and the model built from it, two-stage validation, transcripts and replay, and the base class an adapter completes. Name What it is serialize_observation The observation allowlist, the fields a framework is shown. The framework gets this payload and never the Observation itself. By default the Observation's .engine is the read-only market view, which serves none of the hidden state. Under trusted_agents=True it is the live engine, and through it fair value, the eleven-way attribution of every move, each company's mispricing and the macro path the run has not reached yet. The payload is the same allowlist either way. decision_schema The one definition of a valid decision. decision_model builds the Pydantic model from it. parse_decision The first validation stage. It checks the answer against the schema. An unknown top-level key, or no actions list, refuses the whole decision with DecisionError. A bad action, such as one with an unknown field, a negative quantity or a limit order with no price, is refused on its own: it goes into the returned Decision's refused list with its reason, and the other actions stand. orders_from The second validation stage. It checks each symbol against the listed universe and each size against the participation cap, the largest fraction of average daily volume one order may take, and turns the actions into orders: a share count for a market order, a tf.Limit for an action with a limit_price, and a tf.Cancel for CANCEL. An order over the cap is clipped, and the clip is recorded. An unlisted symbol goes into the refused list it is given, and raises MarketRefusalError when it is given none. Transcript Record and replay. Each exchange is keyed by a digest, a hash of the exact input sent to the framework, and meta names the market the recording was made in. run_sync The one supported bridge from sync code to async code. It runs a coroutine on another thread, so context variables the caller set do not carry into it. FrameworkAdapter The base class an adapter completes. ReplayMixin adds the record-and-replay path. Every adapter shares the constructor arguments info, every, fundamentals, max_participation and arm. The module exports seven constants, OBSERVABLE_MACRO , SIDES , ORDER_TYPES , HISTORY_STEPS , MAX_PARTICIPATION , DECISION_SCHEMA_VERSION ("2" since 0.8.5) and OBSERVATION_SCHEMA_VERSION ("1"). Both versions are frozen for the 0.8.x line. The observation payload lists the payload, and Bringing an LLM agent lists the decision. It also exports three classes, Action , Decision and AdapterInfo , and three functions, require , digest and replay_response . It also exports the preset check the built-in adapters run, so an adapter you write yourself can run it too. The check is preset_of(obs) , stamp_preset(recorder, obs) , refuse_a_changed_preset(transcript, preset) , stamp_artefact(meta) and the constant PRESET_VECTOR_KEY . Error behavior Every error below is a subclass of tradefloor.ValidationError . A bad answer from the agent raises DecisionError , and a failure in the framework raises FrameworkError . The two are kept apart so that a weak agent is not mistaken for an unreliable network. Name Meaning IntegrationError The root of this error family. MissingDependencyError Also an ImportError. require() raises it. FrameworkError The framework call itself failed, for example with a timeout, a transport failure or a budget stop. DecisionError The agent's output could not be read as a decision. The error names the step and the day. MarketRefusalError A DecisionError. The decision is well formed, and this market cannot take it. The adapters pass a refused list, so one bad action is refused on its own and the rest of the decision trades. ReplayMiss A DecisionError. The recording has no answer for this input, or it was made in a different market. World re-raises it under on_refusal="skip", so a replay that cannot answer stops the run instead of counting against the agent. No adapter turns a failure into an empty decision. An empty decision scores as trades=0 with an empty error column, which is also how an agent that chose not to trade scores. If failures became empty decisions, an agent that failed and an agent that declined would score the same. Replay and recording A Transcript keys each exchange by a digest of the exact input sent to the framework, never by a step number. If you change the roster, the seed or the cadence, the key has no recorded answer, so the run stops with ReplayMiss and names the step. With a step-number key, a changed experiment would get the answers recorded for the old one, and nothing in the output would show it. Instructions are checked separately, because most adapters do not send them in the keyed input. LangGraphAdapter renders its instructions into that input, so a change stops the run as above. The OpenAI Agents, PydanticAI and FinRobot adapters write a digest of their instructions to meta["instructions_digest"] , and a replay under different instructions raises ValidationError when the adapter is built. CallableAgentAdapter cannot see the prompt inside your function, so it checks only when its AdapterInfo carries an instructions_digest . adapter.record holds one entry per decision, with the digest, the exact input, the raw response, the validated decision, the orders, any participation clips and any refused actions with their reasons. A limit order is held as {"quantity", "limit_price"}. So the whole path from observation to order is in one place, in a replayed run as in a live one. Recorded preset Every price in an observation comes from the preset, the named set of model coefficients the engine runs. So a recording only replays against the preset it was made in, and the recording names that preset. On the first recorded exchange, every adapter writes the running engine's fingerprint to meta["model_preset"] and the full parameter vector to meta["model_preset_vector"] . The fingerprint is a preset name such as pt-v19 or custom-XXXXXXXX . Transcript.save writes recorded_utc . For a transcript that never ran against an engine, it also writes the shipped default as model_preset . finrobot.Transcript.save writes the same two fields. A replay also refuses a different market. replay_response(transcript, key, *, step, day, preset=None) compares the recorded preset with preset before it looks up the digest, and every adapter passes the running engine's preset. A different name raises ReplayMiss , naming both. To fix it, replay on the recorded preset with World(..., model=...) or evaluate(..., model=...) , or record the run again. The same preset name can also carry different values, which happens when a preset is cut again between builds. That also raises ReplayMiss , and the error lists the dials (model parameters) that moved. A recording made before 0.8.0 has no model_preset , and a replay of it warns and goes on. A recording's meta also carries observation_schema_version and decision_schema_version , and a replay of a recording made under another payload version stops before the first lookup and names both versions. A recording made before 0.8.5 carries neither, and its first lookup misses, because the payload it was keyed on changed. Replay it on the release that made it. Generic callable CallableAgentAdapter wraps a plain Python function so that evaluate can run it, and callable_agent(fn, **kwargs) is the convenience constructor. An async function runs through common.run_sync . The function gets the serialized payload, never the Observation. CallableAgentAdapter(fn=None, *, name, info, every=6, mode="live", transcript, recorder, prior, postprocess, ...) . fn may be left out in replay mode, which never calls it. postprocess(raw, payload) runs on whatever fn returned, in a live run and in a replay, and its return is the decision. The transcript records fn 's return, the raw model response, so parsing, risk checks and sizing written in postprocess are exercised by every replay. Code inside fn after the model call is recorded as its output and never runs on replay. Without postprocess , fn 's return is the decision. The replay key is the payload alone. Pass info=AdapterInfo(framework="callable", instructions_digest=digest(PROMPT)) when you record and when you replay. The digest is written to the transcript's meta, and a replay built with a different digest raises ValidationError before the market opens. With no AdapterInfo , a replay under an edited prompt runs to the end on the recorded answers. Replay checks has a worked example. OpenAI Agents SDK adapter OpenAIAgentsAdapter(agent, *, mode="replay", transcript, recorder, model, brief, max_turns=6, tracing=False, run_id, ...) . The convenience constructor openai_agent(agent, **kwargs) defaults to mode="live" . The module also exports payload_of(call) , BRIEF , DISTRIBUTION and EXTRA . The adapter adds two methods to the base, ask(obs, payload) and input_items(payload) . state() adds max_turns and brief_digest , because both change what a decision can be. Compatibility The adapter binds the decision schema as the output type on a copy made with Agent.clone(...) . Your agent is left unchanged, and its instructions, tools, model settings, hooks, handoffs and guardrails all carry over. An agent that already declares its own output_type is refused, and the message says how to opt in. A run that ends on a different agent through a handoff is refused by name, because that agent has an output type of its own. Runtime behavior The adapter calls Runner.run through the shared async bridge. On 0.22.0, Runner.run_sync cannot run inside an existing event loop, so it does not work from a notebook. max_turns caps one decision at six model calls. The SDK's default of ten suits interactive use, and a loop that runs at every cadence step of every experiment arm needs the lower cap. The adapter imports the SDK only inside the method that calls it, so a replayed run needs neither the package nor the time its import takes. Retry and validation behavior On 0.22.0, the SDK makes one model call on a malformed answer and raises ModelBehaviorError . It does not retry on the client side. Binding the decision model puts the side enum, the non-negative quantity and additionalProperties: false into the schema the provider sees, which makes an invalid decision less likely. The binding adds no repair loop, so write your error handling for this adapter as if there were no retries. Error behavior The adapter treats seven SDK exceptions as the agent's own outcome and turns each into a DecisionError : ModelBehaviorError , ModelRefusalError , MaxTurnsExceeded and all four guardrail tripwires. By default the SDK removes model text from its own messages, so its message alone cannot say which decision point failed. A tripped guardrail is not converted to a hold. Every other exception stays a FrameworkError with the exception chain intact. Tracing SDK tracing is on by default and sends traces to OpenAI. This adapter passes tracing_disabled=True on every run it starts, unless you gave tracing=True . It does this per run instead of through the SDK's process-wide switch, so tracing for other code in the process is left alone. Version notes The minimum version is openai-agents>=0.22 , the version the adapter was written against. A minimum at the major version would allow releases the adapter has not been tested against. The package imports as agents and supports Python 3.10 through 3.14. tradefloor needs 3.11 or later, so every Python that runs tradefloor runs the SDK. PydanticAI adapter PydanticAIAdapter(agent, *, deps, mode="live", model, transcript, recorder, instructions, bind_output_type=True, request_limit=8, ...) . The module also exports UsageLimitReached , render(payload) , MANDATE and MANDATE_VERSION . Compatibility You pass in a built agent, and the adapter does not modify it. Its deps_type , tools, RunContext usage, instructions, toolsets and output type keep working, and deps reaches run(deps=...) unchanged. PydanticAI has one dependency slot. Every tool reads it through RunContext.deps , and nothing checks its type at runtime, so an adapter that put its own payload there would hand your tools an object of the wrong type. The adapter puts nothing in it, so a tool cannot query the observation from inside a decision. A tool that needs the day's prices reads them from a holder on your own deps object. Decision mapping The adapter binds the shared decision model as the run's output type, so the side enum, the non-negative share count and the required actions list are in the schema the model sees. The binding applies to that run only, and your agent's own output type is untouched. PydanticAI does not allow an override of the output type on an agent with an @agent.output_validator . In that case the adapter says so and points to bind_output_type=False . With that set, your output type stays and tradefloor still validates what it produces. Retry and validation behavior PydanticAI's own retry loop catches a schema violation and corrects it within the turn. The adapter files UnexpectedModelBehavior as a DecisionError , because an agent that used up its retries without a valid decision answered badly. A run that hits its request limit raises UsageLimitReached , a FrameworkError subclass, so you can catch a deliberate budget stop by name. Runtime behavior Pass an offline model through the adapter's model= argument instead of Agent.override . Agent.override is built on context variables, and those do not carry across the shared async bridge. A test suite can set models.ALLOW_MODEL_REQUESTS = False . A blocked request then raises a plain RuntimeError, so a test that expects a framework exception will miss it. Observability PydanticAI records no traces by default, and the adapter turns nothing on. In the supported version, you set up tracing through PydanticAI's own current APIs, either logfire.configure() with logfire.instrument_pydantic_ai() , or Agent.instrument_all() . 2.36.0 has no instrument= constructor argument. Version notes The minimum version is pydantic-ai-slim>=2.36 . The slim package has the pydantic_ai module without the provider SDKs that the full package adds, and no adapter imports any of those. TestModel and FunctionModel are both in slim. LangGraph adapter LangGraphAdapter(runnable, *, mode="live", transcript, recorder, input_builder, output_parser, instructions, config, thread_id, ...) , with langgraph_agent(runnable, **kwargs) as the convenience constructor. The module also exports default_input_builder , default_output_parser , render , GraphInterruptedError , INSTRUCTIONS , DEFAULT_INPUT_KEYS and INTERRUPT_KEY . Compatibility The adapter accepts any object with an invoke method and does not check its class. Runnable is a nominal ABC with no __subclasshook__ , so an isinstance check would reject a plain object with a working invoke , and a deterministic test double is that kind of object. An object with only ainvoke runs through common.run_sync , and an uncompiled StateGraph is refused by name. Input mapping The default input carries both input shapes at once, observation for a structured graph and messages for the MessagesState shape. A key the graph's state schema does not declare is dropped before any node runs, so one default works for both. A TypedDict state gets no schema check at the graph boundary, so a mismatch shows up as an IndexError or a bare KeyError inside your own node. Use input_builder when the graph expects a different state schema. Output mapping A graph returns its whole state, so the adapter has to take the decision out of it. The default parser reads a Decision, an interrupted state, a state carrying actions , a state carrying decision , or the MessagesState shape. It refuses anything else by name. Interrupts An interrupt raises GraphInterruptedError , a DecisionError subclass, which names the question that went unanswered. An interrupt means pause now and resume later. A market loop has nowhere to resume into, because the order book, the macro path and the variance process move on as soon as act returns. On 1.2.11, a GraphInterrupt never escapes invoke and arrives inside the state as __interrupt__ . The parser checks for it before it looks for decision , because a checkpointed thread can carry a decision written on an earlier step. Run a human-in-the-loop graph separately, and give tradefloor a graph that decides. Observability The adapter puts the run's identity in config["metadata"] and adds a "tradefloor" tag, merged into any RunnableConfig you pass. It uses metadata because a node receives tags , metadata and recursion_limit from the invoking config, while the tracer consumes run_name . LangSmith tracing stays off unless one of its environment variables is true . Every exported run carries tradefloor_run_id . The adapter also stamps tradefloor_arm , tradefloor_day , tradefloor_step and tradefloor_decision_schema on each decision. Turning tracing on sends the rendered observation to LangSmith. Version notes The minimum version is langgraph>=1.2 . That one requirement is enough, because langchain-core is a hard dependency of it. LangGraph needs Python 3.10 or later, so with tradefloor it runs on 3.11 and later. create_react_agent is deprecated in LangGraph 1.x and points at a package this extra does not install. A plain StateGraph has no deprecated import, and the adapter treats both the same way. FinRobot adapter The FinRobot integration is older than the shared layer, which was built from it. Its DecisionError inherits from common.DecisionError instead of directly from ValidationError . It is still importable, it is still a ValidationError , and it can also be caught as the shared error. It also runs the shared layer's preset check. A FinRobot recording names its market in meta , and a replay against a different one raises ReplayMiss . The finrobot extra requires exactly Python 3.11, because FinRobot declares >=3.10, <3.12 and tradefloor needs >=3.11 . Replaying a recorded run needs none of it. ====================================================================== # Gym environment https://docs.tradefloor.dev/rl-environment.html A Gymnasium environment where a policy's own orders move prices, so it pays for the size it trades. Actions are target weights and reward includes the impact. ====================================================================== API REFERENCE/ GYM ENVIRONMENT Gym environment tradefloor.gym.TradingEnv is a Gymnasium environment, the standard Python interface for reinforcement learning. The policy's orders trade in the market's book and move its prices, and each seed is a new market. TradingEnv TradingEnv( *, universe: Sequence[Instrument], seed: int = 0, macro: Macro | None = None, days: int = 5, steps_per_day: int = 6, ticks_per_step: int = 65, cash: float = 1_000_000.0, max_leverage: float | None = 2.0, start: tuple[int, int, int] = (9, 30, 3), model: str | ModelParams | None = None, trusted_agents: bool = False, ) Argument Default Meaning universe required The roster. Its order is part of the market. seed 0 The first episode's market, and the seed of the generator that draws later episodes' seeds. macro None The economy on day 0. None uses Macro()'s defaults. days 5 Trading days in an episode. steps_per_day 6 Steps a day. An episode is days * steps_per_day steps. ticks_per_step 65 Simulated minutes the market runs after each action. cash 1000000.0 Starting cash. max_leverage 2.0 Cap on gross exposure as a multiple of net worth. None removes it. start (9, 30, 3) Hour, minute and day of the week the first session opens on. model None A preset name or a ModelParams, the same for every episode. None is the default preset, pt-v20. trusted_agents False True makes env.engine and env.portfolio the live objects. By default they are read-only views. It needs numpy and gymnasium, from pip install "tradefloor[rl]". It passes Gymnasium's env_checker, which warns only that the observation space is unbounded. Spaces and reward For a roster of n names: Space Holds Observation Box(-inf, inf, (2n + 1,), float64) Each name's log return since the previous step, each holding as a fraction of net worth, then cash as a fraction of net worth. Action Box(-1, 1, (n,), float64) A target weight per name, as a fraction of net worth. Negative is short. The observation space is unbounded because no finite bound holds: a step's return is limited only by the circuit breaker on each of its ticks, and cash goes negative when the book is levered. An action outside [-1, 1] is clipped. When the absolute weights add up to more than the leverage cap allows, every weight is scaled by one factor to a gross exposure of max_leverage / (1 + 0.01 * max_leverage), 1.96 under the default cap of 2, and info["scaled"] is True. The step sells what it shrinks before it buys what it grows. A trade the book or the cap refuses is counted in info["rejected"], and the episode goes on. The reward is the step's change in net worth, in dollars: net worth after the step's session less net worth before it. It is measured after the market has moved, so it includes the cost of the policy's own trading. A step's fills reach the market once, on the first tick of the session that follows, so rewards for a policy that trades differ from 0.8.1 and earlier. An episode is terminated when net worth reaches zero and truncated after its last step. info key Returned by Meaning seed reset The seed the episode ran. reset(seed=info["seed"]) replays it. model_fingerprint reset The preset's name, or custom- and a hash for changed coefficients. trusted reset True, present only under trusted_agents=True. net_worth, cash step Net worth and cash after the step. leverage step Gross exposure as a multiple of net worth. rejected step Trades refused this step. scaled step Whether the action was scaled down to fit the leverage cap. step step Steps taken in the episode. An episode import numpy as np import tradefloor as tf from tradefloor.gym import TradingEnv universe = tf.Universe.random(5, seed=101) env = TradingEnv(universe=universe, seed=42, days=5) obs, info = env.reset() print(obs.shape, env.action_space.shape, info["seed"]) pnl = 0.0 done = False while not done: action = np.full(5, 0.6) # 60% of net worth in each name: 3x gross obs, reward, terminated, truncated, info = env.step(action) pnl += reward done = terminated or truncated print(info["step"], info["scaled"], round(info["leverage"], 2), round(pnl)) (11,) (5,) 42 30 True 1.97 38961 The action asks for 3 times net worth, so every step scales it to fit the cap of 2. Seeds and episodes The first reset() runs the constructor's seed, and reset(seed=n) runs seed n. Each later reset() without a seed draws a new seed below 2**32 from the environment's generator, so a loop of 1,000 resets meets 1,000 markets, the same 1,000 on every run. To see how much of a score came from the market, hold the universe fixed and change the seed. The package version, the preset, the universe's fingerprint and the seed fix an episode, so someone else can replay the episodes a policy trained on. Limits of a trained policy A policy that learns the model Prices come from a known model, and a policy is very good at finding that model's structure. A high score can mean the policy found the herding term, a property of the model that no real market has. Compare the trained policy with buy-and-hold across many seeds, and before you make a claim about a real market, read How it is measured, which shows where the model matches real markets. Weaker volatility memory than real markets A volatile period fades faster in the model than in a real market. On pt-v20, volatility clustering, the tendency of volatile days to follow volatile days, is below real at every lag: about a quarter of the real strength one day apart and a sixth twenty days apart. A policy that sizes its trades on a volatility estimate of one month or longer learns less memory than real markets have. ====================================================================== # Model parameters https://docs.tradefloor.dev/parameters.html Reference for ModelParams and every entry it exposes: the settable names, their types, their values under the shipped preset, and what each one controls. ====================================================================== API REFERENCE/ MODEL PARAMETERS Model parameters Reference for ModelParams , the set of model coefficients an Engine runs, and for every entry in it. A preset is a named set of coefficients that ships with the package. The values shown are for preset pt-v20 , the shipped default. The release notes list what each preset changed, and tf.model_preset() prints the default preset's dictionary at run time. 243 entries in to_dict() 214 settable 29 not settable 19 shipped presets ModelParams A set of coefficients. Build one from a shipped preset, optionally overriding settable entries, and pass it to Engine(model=...) . Instances are immutable. There is no setter, and every constructor returns a new object. ModelParams.from_preset Build a coefficient set from a shipped preset, optionally overriding settable entries. ModelParams.from_preset( name: str = "pt-v20", **overrides: float, ) -> ModelParams Argument Type Default Meaning name str pt-v20 A shipped preset name, pt-v1 to pt-v20. If you leave it out, you get the shipped default, the same one Engine runs. **overrides float - Settable parameter names, from the table below. Any other name is refused. Returns ModelParams. Raises ValidationError on an unknown preset name, an unknown parameter name, a compile-time entry, or an override that breaks an identity the preset claims. An identity is a fixed relationship between parameters, and only pt-v19 claims any: garch_beta is 0.9416 - garch_alpha - garch_gamma / 2, vix_return_level_exponent is vix_return_exponent - 1, and vix_target_shock_cap is the ceiling the VIX return dials imply. So from_preset("pt-v19", garch_alpha=0.07) raises unless garch_beta moves to match. The message names the identity, both values and the tolerance. ModelParams.from_preset_unchecked Works like from_preset but skips the preset's identity checks. It is for measuring a derived parameter moved away from the value its identity gives, and any number it produces should state that the checks were skipped. ModelParams.from_preset_unchecked( name: str = "pt-v20", **overrides: float, ) -> ModelParams Argument Type Default Meaning name str pt-v20 Same as from_preset. **overrides float - Same as from_preset. Returns ModelParams, the same immutable type, bit for bit what from_preset would build. An override fingerprints as custom- as usual. Raises ValidationError on everything from_preset refuses except a broken identity. The rules that apply to every preset are still checked. ModelParams.identity_breaks Checks a set of coefficients against the identities a preset claims. ModelParams.identity_breaks( params: ModelParams, preset: str = "pt-v20", ) -> list[dict[str, Any]] Argument Type Default Meaning params ModelParams required The coefficients to check. preset str pt-v20 The preset whose identities to check. Each preset has its own, so pass the one you mean. Returns list[dict], one per identity that does not hold, with dial, identity, expected, actual, tolerance and claimed_by, or an empty list when all hold. Raises Nothing for a shipped preset name. ModelParams.from_dict Rebuild a coefficient set from a to_dict() mapping. ModelParams.from_dict( values: dict[str, float | str], ) -> ModelParams Argument Type Default Meaning values dict[str, float | str] required A mapping as to_dict() returns it, including the compile-time entries and name. Returns ModelParams. A round trip through to_dict() recovers the preset fingerprint when no value was changed. from_dict does not check preset identities, so run identity_breaks on a rebuilt set if you need that check. Raises ValidationError when a required entry is missing or a value is the wrong type. ModelParams.settable The names accepted as overrides. ModelParams.settable() -> list[str] Argument Type Default Meaning Returns list[str], sorted. The same set the tables below document. Raises Nothing. ModelParams.to_dict Every entry the model carries, settable or not. params.to_dict() -> dict[str, float | str] Argument Type Default Meaning Returns dict[str, float | str]. Every value is a float except name, which is the fingerprint as a str. Raises Nothing. ModelParams.fingerprint The model fingerprint. It is the preset name for an unmodified preset, or custom- after any override. params.fingerprint -> str Argument Type Default Meaning Returns str. A shipped preset's name when nothing was overridden, otherwise 'custom-' and an eight-hex-digit hash of the whole set. Raises Nothing. Parameter classes Class Count Runtime override In to_dict() Fingerprinted settable 214 Yes, as a keyword to from_preset() Yes Yes derived 2 No, recomputed at construction Yes Yes compile-time 27 No, refused by name Yes Yes hidden 2 No No Yes Two struct fields, breaker_down and breaker_up , are model parameters that to_dict() does not return. You cannot read them back from a model, and the tables below do not list them. Fingerprint behavior A model built from a preset has that preset's name as its fingerprint. If you override any settable entry, the name becomes custom- plus a hash of the whole coefficient set, so a run with changed coefficients cannot be cited as the preset. Every preset shows what each shipped set measured, and RunManifest shows what a manifest records. import tradefloor as tf tf.ModelParams.from_preset("pt-v14").fingerprint # 'pt-v14' tf.ModelParams.from_preset( "pt-v14", momentum_theta=0.05).fingerprint # 'custom-b0a8c73d' The hash covers every entry to_dict() returns, so a release that adds entries changes the hash of the same override: the example gives custom-b0a8c73d on 0.8.5 and 0.8.6. A shipped preset keeps its name across releases. When you cite a custom set of coefficients, give the package version with it, or the full to_dict() . Preset records The measurements taken for each shipped preset, read from JSON files included in the package. tf.preset_names() lists the presets the engine knows, and tf.preset_records() lists the ones with a record, which is all of them. tf.preset_record(name: str | None = None) -> dict[str, Any] tf.preset_records() -> list[str] name=None reads the shipped default, pt-v20. An unknown name raises LookupError , and the message lists the names that exist. The record is a plain dict with these keys. Key What it holds preset, fingerprint The preset's name. A record exists for every shipped preset. coefficients, coefficient_digest The coefficients that were measured, and their hash, so you can check that a record describes the coefficients your build runs. default_since The release the preset became the default in, or None. measured The release, commit, roster, seeds and band tables the run used. panel_252, panel_504 The fixed-roster panel at one and two trading years: the median of each statistic over thirty seeds. From pt-v19 it includes crisis_sector_dispersion, the median over the seeds that read it. in_band, misses, unreadable Per cell, the number of rows inside their band (the range real markets show for that statistic), the rows outside their band by name, and the rows the band table has no band for. A row in unreadable counts as neither a pass nor a miss and is left out of the count. absent, dispersion pt-v19 and pt-v20 only. The rows a cell measured but could not read, and how many seeds produced a reading of crisis_sector_dispersion. That statistic needs crisis days, which a calm run may not have. level_protocol The level and crisis rows, measured on a new roster drawn for each seed. pt-v18 to pt-v20. mechanism_252, mechanism_heldout_seeds Whether each row shows its mechanism, tested with a sign test against a reading with the mechanism switched off. structure_252, structure_heldout_seeds, structure_rise The VIX persistence check at one year, and whether persistence rises from one year to two as it does in real market data. crisis_lever Annualized volatility with VIX held at 65, divided by annualized volatility with VIX held at 5, next to the real ratio of 6.16. long_run pt-v19 and pt-v20 only. The long-run rows: the 40 registered for pt-v20, which it passes, and the 17 of pt-v19's record, of which it meets 15. Its keys are criteria, verdict, passed, of, rows and measured. Each entry in rows has id, words, value, real, rule and pass, and measured names the commit, the seeds and the years. tf.preset_record()["long_run"]["rows"] returns the long-run rows with the real value beside each reading. Settable parameters Every name accepted as a keyword override, grouped by the part of the model it belongs to, with the file that implements that part. Every settable and derived entry is a Rust f64 , a float from Python. The seven agent-facing book entries are 0.0 on every preset before pt-v20, which leaves the book there as it was, and pt-v20 sets all seven. They change what Engine.submit and Portfolio.execute meet and nothing in a market no agent trades. Engine and data says what each one does to an order. The agent-facing book agent_book.rs, engine.rs  ·  7 parameters Name pt-v20 Description book_depth_coefficient 0.75 Coefficient Y of the square-root law that sets the latent depth behind the maker's ladder. book_depth_exponent 0.5 The exponent delta of the latent depth's price-for-size law. book_depth_reach 1 How far the latent depth reaches, in multiples of the name's average daily volume per side. book_refill_half_life 27 Half-life in ticks at which consumed LATENT depth refills. book_resting 1 Switch for whether an agent's unfilled limit order rests in the book: 1.0 on, 0.0 off. book_shared 1 Switch for whether agents' orders consume the book they share: 1.0 on, 0.0 off. fill_impact_coefficient 0.314 Permanent impact of an agent's fills, linear in size: gamma in ds = gamma * sigma * (bought - sold) / V, applied to the name's mispricing s once, on the first tick after the fills. Market-factor variance process market/factor_vol.rs  ·  67 parameters Name pt-v20 Description buyback_payout_share 0.75 The share of earnings a company returns as net buybacks, paid as a yield on the price. cascade_gain 0.1 A scale on the whole forced-flow term: the short squeeze and both stop ladders (cascade_symmetry). cascade_symmetry 1 How much of the direction in the stop-cascade ladders is removed, from 0.0 (the original asymmetric ladders) to 1.0 (mirror images). cycle_hazard_per_month 1 Reads the business cycle's hazard as a rate per month: at 1.0 the monthly transition probability is divided by 30 before the daily draw. cycle_stationary_opening 1 Switch (0.0 or 1.0) that draws the day-zero cycle phase and its age from the cycle's own stationary law, instead of opening every run at the start of an expansion. cycle_us_calibration 1 Switch for the business-cycle phase table derived from NBER and BEA data (economy::state::us_phase_characteristics). earnings_nominal_growth 1 How much of nominal output growth the valuation's earnings carry, from 0.0 (earnings fixed at construction) to 1.0 (the earnings share of nominal output held constant). fair_value_book_floor 0 Switch that applies the loss-maker book floor to profitable companies too, making fair value continuous at zero earnings. fed_liftoff_rule 1 Switch for a lift-off branch in the central bank's rate ladder. garch_cascade_components 0 How many components a name's variance cascade carries. garch_cascade_ratio 3 Half-life spacing between cascade components: component k has a half-life ratio^k times component 0's. garch_cascade_weight 1 How much of the variance comes from the cascade rather than from the single-component process. jump_mean_compensated 1 How much of the drift the market jump's mean carries is given back, from 0.0 (none) to 1.0, which subtracts the compensator and makes the jump a martingale. macro_burn_in_days 755 Days the economy is advanced alone, before day zero, so a run opens on settled macro fields. macro_calendar_days_per_year 252 Economy steps per macro year for the rest of the macro calendar: months, quarters, the seasonal year and central-bank meetings. macro_compound_days_per_year 252 Economy steps per year used to compound annual GDP and CPI growth. market_beta_down_asym 0.025 Extra transmission of the market factor on a down tick: every name receives beta * factor * (1 + this). market_beta_down_asym_lag 0.46 The lagged downside transmission: on the session after a down day, every name receives beta * factor * (1 + this) whatever the tick's own sign. market_beta_down_asym_lag_live 1 Where the lagged wire's down-day condition is sampled: at the open (0.0), live through the session (1.0), or with the sign reversed as a diagnostic (2.0). market_beta_down_asym_lag_recentre 0 How much of the extra mean that the lagged down-day wire adds to the tilt is given back, from 0.0 (none) to 1.0 (all of it). market_beta_down_asym_recentre 1 How much of the mean that market_beta_down_asym injects is given back, from 0.0 (none) to 1.0 (all of it). market_burn_in_sessions 0 Sessions of warm-up given to the market factor's variance components before session one. market_idio_down_suppress 0 Shrinks a name's idiosyncratic shock on a down tick of the market factor and inflates it on an up tick, so same-day down-market co-movement rises while the unconditional variance is held exactly. market_pe_buybacks 1 Switch that includes buybacks in the earnings behind market_pe. market_vol_alpha 0.0066 The market factor's own GARCH reaction term: how sharply market-wide variance responds to the last market-wide shock. market_vol_alpha_excursion 0 How far the common factor's shock share moves with the factor's own variance excursion. market_vol_beta 0.8946 The market factor's variance persistence. market_vol_ceiling_multiple 32 Cap on the market factor's variance, as a multiple of its calm level. market_vol_floor_multiple 0.05 Floor on the market factor's variance, as a multiple of its calm level. market_vol_gamma 0.1556 GJR leverage on the market factor's variance update: the extra weight a down day's squared shock gets. market_vol_level_persistence 0 Session-to-session persistence of a slow multiplier on the market factor's variance target, in (0, 1). market_vol_level_sigma 0 Per-session innovation of the slow variance level, in log units. market_vol_slow_gain 0.05 How much of each day's variance surprise the slow component takes up. market_vol_slow_persistence 0.9913 Daily persistence of the slow component of the market factor's variance (Engle-Lee style). market_vol_slow_weight 0.35 Weight of the slow variance component in the market factor's two-component mixture, from 0.0 (single component) to 1.0. market_vol_vix_anchor 15.9843 VIX level at which a coupled target equals the baseline variance. market_vol_vix_coupling 0.95405 How far the market factor's variance target follows the VIX, from 0.0 (a fixed target) to 1.0 (fully proportional to the VIX response, the squared VIX ratio at the default market_vol_vix_exponent). market_vol_vix_exponent 4 Exponent on the market variance target's VIX ratio. market_vol_vix_exponent_below 2.5 Exponent on the market variance target's VIX ratio when that ratio is below one, where the VIX sits under the ratio's denominator (the derived anchor under the level form, the read-back under the excursion form). market_vol_vix_smooth 0 Days of EMA smoothing on the VIX that the market variance target reads. neutral_discount_rate 0.0482 The corporate bond yield at which the target multiple sits exactly on its sector anchor, as a fraction (0.04 is 4 per cent). oil_opec_symmetry 1 Removes the direction from the OPEC production rule while keeping its size. oil_seasonality_target 1 Where oil's seasonal shape acts: on the price level itself (0.0) or on the price the process reverts toward (1.0). oil_supply_response 1 How much of oil demand is answered by supply on the daily step, from 0.0 (none) to 1.0 (supply equals demand in expectation). phase_target_range_draw 0 Whether a phase's growth target is drawn from its declared range (1.0) or fixed at the range's midpoint (0.0). trough_growth_floor 0 Raises the bottom of the trough phase's GDP growth range from -1.0 (at 0.0) to 0.0 (at 1.0), moving proportionally in between; the top stays at 0.5. vix_anchor_centre 0.1515 Log offset below the identity's derived anchor that the anchor weight pulls toward: the blend and the memory's reference use L * anchor * exp(-c). vix_anchor_memory 0.0555556 Per-session rate at which the anchor's memory of the read-back updates. vix_anchor_reversion 0 Per-session rate at which the VIX reverts toward the identity's anchor, as a share of the distance in [0, 1). vix_anchor_weight 0.375 Share of the VIX's target taken by the identity's anchor, as a geometric weight in [0, 1). vix_anchor_weight_level 1 How the anchor weight rises with the VIX's level, as an exponent. vix_anchor_weight_level_below 0 Switch that also runs the level law below the knee, where it lowers the anchor weight toward zero. vix_anchor_weight_level_cap 2.2159 Multiple of the knee above which the level law stops raising the anchor weight. vix_anchor_weight_level_knee 0.3888 Where the level law starts raising the anchor weight, as a log offset below L * anchor. vix_anchor_weight_level_knee_fixed 0 Switch that fixes the level law's knee to the anchor alone, without the slow regime level: K = anchor exp(-k) in place of K = L anchor exp(-k). vix_level_loop_gain 1.79 The VIX loop's own gain on the slow VIX level, which the level's innovation is divided by. vix_level_persistence 0.9979 Session-to-session persistence of the VIX's own slow log-level, a lognormal AR(1) multiplier on the VIX target under vix_level_identity. vix_level_sigma 0.0181 Per-session innovation of the VIX's own slow log-level, in log units. volume_idio_persistence 0 Daily persistence of a per-name volume state, so each name has busy and quiet spells of its own. volume_idio_sigma 0 Innovation size of the per-name volume state. volume_idio_variance_gain 0.2 Gain that makes a name's volume follow its own conditional variance, so a name trades more when its own volatility is high. volume_move_cap 12 Where the volume response to a move saturates, in units of one percent. volume_move_floor 0.6 Base volume multiplier for a name on a day it does not move at all. volume_move_jump_share 1 How much of a jump's share of the day's move the volume scale counts, from 0.0 (none) to 1.0. volume_move_noise 0.2 Amplitude of the return-unrelated noise in a name's daily volume. volume_move_response 0.6 How much more a name trades per one percent it has moved today. volume_variance_gain 0.0284038 How strongly realized volume tracks the market factor's variance. Factor structure market/tick.rs, market/factors.rs  ·  63 parameters Name pt-v20 Description buyback_yield_cap 0.15 A ceiling on the annual buyback yield buyback_payout_share * eps / price that the buyback term compounds over the elapsed years. closing_auction 1 Whether the session closes with a cross at the model price, a switch. corporate_yield_daily 1 Whether the corporate yield moves between central-bank meetings, a switch. crash_amplifier_conditional_sigma 1 Switch for the sigma the crash amplifier measures a shock in: 0.0 uses the baseline constant, any nonzero value the tick's own conditional sigma. crash_amplifier_slope 0.2 Extra market loading the crash amplifier adds per baseline sigma of shock beyond crash_amplifier_threshold. crash_amplifier_threshold 2 Market-shock size, in baseline sigmas, above which the crash amplifier fires. crisis_blend_cap 0.98 Ceiling of the crisis correlation blend. crisis_blend_gain 0 How hard a crisis loads every name onto the market factor, as a multiplier on the crisis spike. crisis_blend_ramp 1.4 VIX points past CRISIS_VIX_THRESHOLD for the sector-to-market crisis blend to reach 1.0 before its cap. crisis_blend_source 1 Where the crisis correlation injection comes from. crisis_blend_variance_damp 0 How far the crisis blend's market injection is decoupled from the market factor's own magnitude, from 0.0 (fully coupled) to 1.0 (fully decoupled). cycle_publication_lag 252 Sessions between a turn of the business cycle and its publication. earnings_anticipation_half_life 126 Half-life, in sessions, of the discount the valuation puts on the earnings cycle's expected path. earnings_cycle_depth 0.2 The aggregate earnings cycle's depth: the log level every company's earnings are pulled toward in a contraction or a trough, beyond what nominal output alone gives them. earnings_cycle_half_life 60 Half-life in sessions of the pull toward the phase's level, 60 on every shipped preset. earnings_cycle_sigma 0 The daily sd of the earnings level's own noise; 0.0 is the phase path alone and takes no draw. earnings_cycle_upside 0.09 The earnings cycle's upside share: every phase other than a contraction or a trough pulls earnings toward +depth * upside, which centers the level over a cycle so the cycle moves earnings around the nominal-output path without shifting it. endogenous_news_intensity 0.05 Daily probability that a company generates its own news event. endogenous_news_sigma 0.01751 Standard deviation of an endogenous news event's price impact, in the units NewsEvent::price_impact carries. fair_value_market_linear 1 Which part of a market shock fair_value_market_share makes permanent, a switch. fair_value_market_share 1 The share of each market-wide shock that moves fair value for good, in [0, 1]: the name's loading on the market factor's draw, market-wide news and the market jump. fair_value_market_vol_cap 1.5 A ceiling, in multiples of market_factor_sigma, on the market volatility whose shocks fair_value_market_share makes permanent, in [0, 32]. fair_value_news_share 1 The share of each idiosyncratic shock that moves the name's fair value for good instead of its mispricing, in [0, 1]. fair_value_vix_discount 0.35 Volatility feedback: a discount on every name's fair value while the VIX is above fair_value_vix_knee, exp(-this * beta * ln(vix / knee)), in [0, 1]. fair_value_vix_half_life 5 Half-life, in sessions, of the VIX exposure the volatility-feedback discount reads. fair_value_vix_knee 40 The VIX level, in points, above which fair_value_vix_discount applies. fear_greed_published_inputs 1 Switch that makes the fear/greed index read the business cycle and GDP growth as published instead of as they are. flight_to_quality_day 1 Which return the flight to quality reads, a switch. flight_to_quality_gain 0.008 The flight to quality's size: percentage points of 10-year yield per percent of index return, down with the market when inflation is under 3 percent and up when it is over 4. gdp_publication_lag 21 Sessions between the end of a quarter and the publication of its GDP growth, as the BEA's advance estimate comes about a month after the quarter. idio_sigma_beta_exponent 0 How strongly a name's idiosyncratic volatility follows its market beta, as an exponent. idio_sigma_scale 0.512598 Multiplier on every name's idiosyncratic GARCH sigma. inflation_ceiling 6 The hard ceiling on endogenous inflation, in percent. inflation_floor -1 The hard floor on endogenous inflation, in percent. inflation_reversion 0.55 How fast endogenous inflation reverts toward its 2% target each month, as a fraction of the gap. informed_flow_fraction 0.35 Share of order-flow impact that is permanent (information), from 0 to 1. macro_publication_repricing 1 Whether a name's price takes the change the close's macro step makes to its fair value at the moment the step is published, a switch. market_factor_sigma 0.00645407 Baseline daily sigma of the shared market factor. market_vol_vix_excursion 0 Switch for the VIX reading the market factor's variance target uses: 0.0 reads the VIX level against a fixed anchor, any nonzero value reads only the excursion above the level the index's own conditional variance implies. news_absorption_drift_half_life 42 The half-life in ticks of the post-news drift part, h_d in news_absorption_half_life's profile. news_absorption_drift_share 0.12 The share of an endogenous news event's move that arrives as post-news drift, after the fast part: d in news_absorption_half_life's profile. news_absorption_half_life 0.6 How fast the market prices an endogenous news event, as the half-life in ticks (minutes) of the fast part of its move. news_market_weight 0.3 Weight of market-wide news on every name. news_peer_vix_coupling 8 How much harder news transfers to a peer in a crisis. news_peer_weight 0.05 Weight of one company's good news on its sector peers. news_peer_weight_down 0.05 Weight of one company's bad news on its sector peers. news_quote_revision 1 Whether the market maker re-quotes on public news, a switch. news_sector_weight 0.5 Weight on sector-wide news, an event tagged with a sector and no company. opening_market_sigma 0.001 The sd of the market's common opening mispricing, the index's own premium over fair value on day zero. opening_mispricing_sigma 0.016 The cross-sectional sd of each name's opening mispricing. order_flow_coefficient 50 Order-flow impact coefficient: the scale of the price move that injected order flow causes, before informed_flow_fraction splits it. order_flow_impact_law 0 Which participation law the order-flow impact multiplier follows, a switch. qe_pe_gain 0 Gain on the QE valuation channel, where the target P/E takes 1 + qe_pe_gain * qe_pe_boost. qe_pe_stock_gain 0 Gain on the QE stock channel: the target P/E takes + qe_pe_stock_gain * ln(qe_assets_ratio), concave in the level of holdings and zero at the neutral baseline. quote_model_weight 1 Where the market maker centers its book, as a weight in [0, 1] on the model price. rate_pe_sensitivity 3 P/E compression per unit of discount rate above neutral, times a name's growth duration: the target multiple's rate adjustment is 1 - (yield - neutral) * rate_pe_sensitivity * duration. sector_factor_sigma 0.00858305 Daily sigma of each shared sector factor, loaded at sector_loading by every member of the sector (market/factors.rs). sector_loading 0.6 How hard a name loads on its own sector's factor, as a multiplier on the sector draw. sector_loading_beta_slope 0.7 How much a name's sector loading follows its market beta. sector_vix_coupling 1 How much the sector draw's variance follows VIX, on the same (VIX / anchor)^2 target the market factor's variance uses (factor_vol.rs). treasury_10y_noise 0.038 The 10-year Treasury yield's daily noise, in percentage points. treasury_2y_noise 0.022 The 2-year Treasury yield's own daily noise, in percentage points. unemployment_adjustment_half_life 84 The half-life, in sessions, of unemployment's response to its cyclical drivers. Crisis gates economy/daily.rs, market/tick.rs, engine.rs  ·  38 parameters Name pt-v20 Description crisis_epicentre_end_sessions 21 How many consecutive sessions under crisis_vix_threshold end a crisis episode. crisis_epicentre_extra 1.93 How much more volatile the crisis epicentre sector's names are than the other sectors' at the same VIX, as a ratio of total volatility. crisis_vix_threshold 30.8833 VIX level at which crisis behavior begins. daily_credit_floor_gain 1 How strongly the credit spread floors are re-applied on every daily step, from 0.0 (off) to 1.0 (both floors in full). forced_flow_beta_exponent 0 How unevenly forced selling lands across names, as an exponent on beta. forced_flow_gain 0 Common forced-selling flow in stress: a log-shock per VIX point above forced_flow_threshold, per day, applied identically to every name. forced_flow_replenish 0 Fraction of the spent forced-selling budget recovered on each day the VIX is at or below forced_flow_threshold (deleveraging capacity rebuilds in calm). forced_flow_reservoir 0 Total forced-selling budget, in VIX-point-days. forced_flow_threshold 40 VIX level, in points, above which forced flow applies. jump_idio_excitation 0 Self-excitation of a name's idiosyncratic jumps: after a jump the name's arrival rate is lambda (1 + h) with h' = decay h + this. jump_idio_excitation_decay 0 The excitation's daily decay, 0.72 [0.48, 0.79] (half-life two sessions) by the same measurement. jump_idio_vix_decoupled 0 Switch that takes the VIX-squared scaling (jump_vix_coupling) off the idiosyncratic arrival rate, leaving it on the market jump. sector_vol_alpha 0.067 Shock weight (alpha) of the per-sector GARCH(1,1) variance state. sector_vol_beta 0.837 The per-sector variance state's persistence term, 0.837 by the same measurement, which pt-v19 and pt-v20 ship. usd_crisis_vix_threshold 25.5 The VIX above which the dollar catches a safe-haven bid. vix_ceiling 181.329 Upper bound on the VIX state itself, in points. vix_cycle_amplitude 0.85 How much of the VIX's level comes from the business cycle, from 0.0 (none) to 1.0 (the full per-phase constants). vix_decay_ratio 1 Multiplier on the VIX mean reversion on days the target sits below the current VIX, so fear can decay more slowly than it arrives. vix_innovation_return_sigma 0.0175 The part of the innovation scale that rises with the session return, per percent: see vix_innovation_sigma. vix_innovation_sigma 0 Base scale of the VIX's own daily innovation, as a fraction of its level. vix_jump_intensity 0 Rate of exogenous fear events, per year, each a jump in the VIX level. vix_jump_level_scale 1.7 A fear event's mean size in units of the day's innovation scale (VIX * sqrt(s0^2 + (c r)^2), or VIX when the innovation dials are off), exponential draw. vix_jump_return_intensity 6.199 The part of the fear-event arrival rate, per year per percent of down session, that rises with the session: the daily probability is (vix_jump_intensity + this * max(0, -r)) / 252. vix_jump_scale 0 Mean size of a fear event, in VIX points (exponential draw). vix_level_identity 1 Switch that sets the VIX level from the index's own conditional variance instead of a table of per-phase constants. vix_mean_reversion 0.27 How fast VIX reverts toward its target, as the fraction of the gap closed each day. vix_realised_vol_weight 0.3 Weight of the market's own volatility in the VIX target, from 0.0 to 1.0. vix_return_clamp 15 The index return is clamped to +/- this before it drives the VIX, in the units of the return source. vix_return_exponent 1.4483 Exponent of the VIX's response to a down day's return: 1.0 is linear, above 1.0 is convex. vix_return_exponent_up 0.5433 The up side's own exponent: the spike on an up session is -gain_up * |r|^this * VIX^(-vix_return_level_exponent_up). vix_return_gain 8.83 VIX points added to its target per unit of a down day's index return, before the clamp and cap below. vix_return_gain_up 0.049 The up-day counterpart of vix_return_gain: how far an up day's index return moves the VIX target. vix_return_level_exponent 0.4483 How the down-side fear response falls with the VIX the session opened from: the spike is gain * |r|^p * VIX^(-this). vix_return_level_exponent_up -1 How the up-side response scales with the level. vix_return_source 1 Which index return the VIX reacts to: the session's final minute (0.0, pt-v1 through pt-v8) or the whole day's (1.0, pt-v9 onward), blended in between. vix_target_offset 0 A constant added to the VIX target, in points. vix_target_shock_cap 158.852 Ceiling on the VIX target's whole excursion, in points: the return spike plus the inflation and shock adjustments. vix_variance_premium 0.252 The variance risk premium pi: how far a real VIX sits above the realized volatility of its own index, as a fraction. Mispricing dynamics mispricing.rs, market/tick.rs  ·  6 parameters Name pt-v20 Description crowd_lean_cap 0.02 Bound on the crowd's daily log-price shock. crowd_momentum_gain 0.02 Crowd herding gain per day on yesterday's change in s. crowd_valuation_gain 0.006 Crowd valuation gain per day on s. mispricing_cap 0.9 Hard bound on |s|. mispricing_half_life_days 60 Trading days for half of a mispricing to decay. momentum_theta 0.0185516 Herding: fraction of yesterday's re-rating that continues today. Per-name GJR-GARCH market/garch.rs  ·  11 parameters Name pt-v20 Description garch_alpha 0.0595072 Weight on yesterday's squared shock: how sharply a name's variance reacts to its own last move. garch_beta 0.7905 Weight on yesterday's variance: how long a name's volatility remembers. garch_ceiling_multiple 5 Ceiling on a name's GARCH variance, as a multiple of the sector's long-run variance. garch_floor_multiple 0.25 Floor on a name's GARCH variance, as a multiple of the sector's long-run variance. garch_gamma 0.183185 GJR leverage-effect asymmetry: the extra weight a negative shock gets in a name's next variance. garch_innovation_commensurate 0 Feeds the per-name GJR-GARCH the name's own noise, in the units the coefficients were fitted in, in place of the whole random_noise column. garch_omega 0.000002 The GJR-GARCH constant: the variance a name reverts toward when neither yesterday's shock nor yesterday's variance pulls it. garch_omega_sector_scaled 0 Switch that scales garch_omega by each sector's base variance in place of one constant for every sector. garch_vix_coupling 0.142196 How much a name's own variance follows the VIX, on the market factor's own target shape. garch_vix_exponent 2 Exponent on the VIX ratio in a name's variance reference. idio_sigma_floor 0.0001 The absolute floor under a name's daily variance in the tick and in the overnight path, in daily variance units. Endogenous jumps engine.rs, applied at the day close  ·  10 parameters Name pt-v20 Description garch_beta_dispersion 0 Cross-sectional spread in volatility persistence, in raw beta units. jump_intensity_idio 0.00688953 Daily probability that a per-name idiosyncratic jump fires. jump_intensity_market 0.0282877 Daily probability that a market-wide jump fires. jump_market_variance_share 0 How much of the market jump's log return joins the day's factor innovation, so the GJR variance update sees a crash day. jump_mean_market -0.00852183 Mean of the market jump in log-return units. jump_momentum_share 0 How much of a jump the herding term is allowed to continue, in [0, 1]. jump_sigma_idio 0.075208 Standard deviation of the idiosyncratic jump, in log-return units. jump_sigma_market 0.00245976 Standard deviation of the market jump, in log-return units. jump_vix_coupling 0.2626 How much a jump's arrival rate follows the VIX. overnight_variance_ratio 0 The variance of the overnight move as a fraction of a session's, per name. Universe memory market/tick.rs, engine.rs  ·  4 parameters Name pt-v20 Description market_vol_slow_vix_damp 0.374 How far the slow variance component's target is decoupled from VIX, in [0, 1]. regime_stress_points 0 Stress the business cycle adds to the correlation blend, in VIX-equivalent points at full intensity (a contraction). universe_stress_decay 0 Daily decay factor of the universe's remembered stress level, which keeps crisis correlation elevated after the VIX falls back. universe_stress_weight 0 How much of the remembered stress reaches the correlation blend. Session guards market/tick.rs  ·  2 parameters Name pt-v20 Description price_breaker_fraction 0.25 Circuit-breaker band as a fraction of the session open (+/-25% shipped). price_hard_cap 50000 Absolute cap on any model price (50,000 shipped). Continuous size effect market/factors.rs  ·  4 parameters Name pt-v20 Description size_effect_exponent 0.15 Exponent of the continuous size effect: (cap / 25B) ^ -exponent. size_effect_smoothness 0 Blend from the four-tier size step toward a continuous power law, in [0, 1]. spread_size_exponent 0.455 Exponent of the continuous spread curve. spread_size_smoothness 0 Blend from the four-tier spread step toward a continuous power law, in [0, 1]. Persistent volume engine.rs close, market/tick.rs phase 3  ·  2 parameters Name pt-v20 Description volume_innovation_sigma 0.21 Standard deviation of the daily log-volume innovation. volume_persistence 0.7 Day-to-day persistence of the shared volume component, in [0, 1). Derived parameters These appear in to_dict() and are covered by the fingerprint, but you cannot override them, because they are computed from other entries when the model is built. Name pt-v20 Description mispricing_phi 0.988514 Daily AR(1) coefficient of the mispricing, derived from mispricing_half_life_days. s_phi_tick 0.99997 Per-tick decay, mispricing_phi^(1/390). Compile-time entries Constants fixed when the engine is compiled. They appear in to_dict() and are covered by the fingerprint. An override is refused by name, because the engine would ignore it and the fingerprint would then record a change the run did not make. Name pt-v20 book_levels 10 daily_shock_cap 0.15 default_sector_anchor_pe 18 fair_value_floor 0.01 fiscal_multiplier 0.3 gold_equilibrium_base 2200 gold_mean_reversion 0.002 growth_duration_scale 2 inflation_target 2 inventory_limit_levels 12 loss_making_price_to_book 1.2 name pt-v20 oil_baseline 75 phillips_curve_coeff 0.2 rate_adjustment_floor 0.5 sector_daily_sigma_consumer_discretionary 0.018 sector_daily_sigma_consumer_staples 0.008 sector_daily_sigma_energy 0.015 sector_daily_sigma_financial_services 0.015 sector_daily_sigma_healthcare 0.018 sector_daily_sigma_industrials 0.015 sector_daily_sigma_materials 0.015 sector_daily_sigma_real_estate 0.008 sector_daily_sigma_technology 0.025 sector_daily_sigma_telecommunications 0.01 sector_daily_sigma_transportation 0.015 sector_daily_sigma_utilities 0.008 See also Every preset lists the shipped coefficient sets and what each measured. Engine and data documents the Engine that runs one, and How it is measured says what the default preset reproduces. ====================================================================== # How prices are made https://docs.tradefloor.dev/how-prices-are-made.html The mechanism behind every simulated price: fair value from earnings and rates, a mispricing that reverts, volatility read off the VIX, and an order book. ====================================================================== MODEL AND VALIDATION/ HOW PRICES ARE MADE How prices are made Every price is a fair value, set by the company's earnings and the economy, times a mispricing that wanders and reverts, traded through an order book. The two layers meet in one number, the model price, and the book turns it into the prints an agent trades against. This page describes the default model, pt-v20. The full specification, every equation with the line of code that implements it, is docs/MODEL.md in the library repository. The two layers the economy, daily at the close business cycle -> growth, unemployment, inflation -> central bank -> policy rate -> 2-year, 10-year, corporate yields -> nominal output, earnings cycle -> the VIX every tick fair value V = sector P/E x earnings x rate term x the company's own level mispricing s = mean reversion + herding + crowd + news + order flow + noise model price = V x exp(s) printed price = model price traded through the maker's book daily at the close volatility = per-name GJR-GARCH, market and sector variance, all scaled by the VIX A session is 390 one-minute ticks. At the open the day's company news is drawn. Each tick draws the market, sector and company shocks, updates every name's mispricing and fair value, and settles the print through the book. At the close the variance processes update, jumps land, and the economy steps: the VIX, the yield curve, the business cycle, the earnings cycle and the central bank. The next session reads the new rates and volatility. Fair value A profitable company is worth its sector's anchor P/E times its earnings, times a rate term that shrinks the multiple when the corporate bond yield is above its neutral level, more for fast-growing companies. A loss-making company is valued at 1.2 times book. Earnings grow with nominal output and an aggregate earnings cycle that falls in a contraction, plus buybacks. So rates reach prices the way they do in a real market. At the neutral rate, 100 basis points on the corporate yield moves a profitable company's fair value by 3% to 5.4%, depending on its growth. Rows R6 and E1 of the registered checks compare the market P/E's response to the 2022 rate path and the fall in earnings around a contraction with real data. Part of every shock is permanent. A company's own news and its sector's shocks move its fair value for good, and so do the market's plain shocks up to a volatility ceiling. What a fear regime adds above that ceiling sits in the mispricing and reverts as the fear passes. This matches evidence that mean reversion in index returns concentrates in turbulent periods (Poterba and Summers 1988; Kim, Nelson and Startz 1991). Above a VIX of 40, fear also marks fair value down until it calms, after French, Schwert and Stambaugh (1987). Mispricing The mispricing is the log gap between the model price and fair value. Each tick it moves by: Term What it does Mean reversion pulls the gap back toward zero, with an effective half-life of about 40 sessions Herding a share of yesterday's re-rating carries on today, set by momentum_theta, 0.0186 on pt-v20 The crowd buys what trades below fair value and chases yesterday's move a little, up to a cap News company, sector and market news as it is priced in Order flow the permanent part of order imbalance, including your own fills Squeezes and cascades forced buying when a rising price meets high short interest, and stop-loss cascades either way Noise the market, sector and company random draws, scaled by beta and the variance state Summed over a session the gap follows a stationary daily AR(2), so it neither trends away nor sits still. Herding is low enough that the lag-one autocorrelation of returns stays inside the real band, so a rule that reads only past prices finds no edge a real market would not give it. Volatility and the VIX Each company's variance is a GJR-GARCH process, so a fall raises volatility more than a rise of the same size. The market factor and each sector have their own variance. All three read the VIX, so a high VIX is a violent market and there is no separate crisis switch. Each day the VIX is computed from the index's own variance. It moves toward the volatility that variance implies, plus a fear response to the day's return, and every variance process reads it back. When it crosses the crisis threshold a crisis episode starts, and with probability 0.6 it starts in financial services, as 2008 did. Rows A1 to A3 of the registered checks replay 2008 and 2020 with the real VIX imposed. The economy The business cycle moves through expansion, peak, contraction, trough and recovery. Growth, unemployment and inflation follow it, and a central bank sets the policy rate at its meetings. The 2-year and 10-year Treasury yields and the corporate yield follow the policy rate, with daily moves of their own. An agent sees the economy as it is published. The business-cycle phase and GDP growth arrive late, as the statistical agencies publish them, and a rate decision is priced at the moment it is announced. The order book A market maker quotes around a blend of the last print and the model price, and the market's own flow trades against it, so the print follows the model price with the noise a real tape has. An agent's order meets three kinds of liquidity: the maker's ladder, other agents' resting orders at their limits, and latent depth behind them priced so that cost grows with the square root of the order's size, as Tóth and colleagues measured in 2011 (row C9 of the registered checks). A limit order waits behind the shares already at its price and fills in parts. Depth an order takes refills with a half-life of 27 ticks, and each agent's net fill leaves a permanent impact on the mispricing on the next tick. The eleven factors engine.truth() books every tick's change in mispricing to eleven named factors that add up to it, with a residual of about 1e-16 from rounding. Factor What it books reversion the pull back toward fair value momentum herding: yesterday's re-rating carrying on crowd_lean the crowd's net flow company_news news as it is priced in order_flow_impact the permanent part of order imbalance, your orders included short_squeeze_effect squeezes and stop-loss cascades random_noise the market, sector and company random draws circuit_breaker the correction when the price leaves the session's band jump the daily jump, recorded on the tick it is first seen overnight the close-to-open move, zero on every shipped preset fair_value_shift the part of a shock that moved fair value instead, with a minus sign An agent's explain(day) answers with one of the first ten, and Agents and evaluation says how the answer is scored. Seeds and streams Each process draws from its own random stream derived from the run's seed: the market, the economy, news, jumps, volume and the crisis epicenter. The market stream's schedule depends only on the roster, never on a price or a preset, so two presets run on the same seed see the same market shocks and differ only in how they respond. A seed can be any 64-bit integer. Every coefficient is listed, with its value on each preset, under Model parameters. ====================================================================== # Why pt-v20 https://docs.tradefloor.dev/why-pt-v20.html What the default model gets right about crashes, rates, earnings and trading costs, why that matters when you test a strategy or an agent, and its limits. ====================================================================== MODEL AND VALIDATION/ WHY PT-V20 Why pt-v20 pt-v20 is the default model in tradefloor 0.8.7. A strategy or an agent tested on it meets crashes, recessions, rate shocks and its own trading costs at about the rates it would meet them in a real market. It passed all 40 checks registered for it before its grade, on seeds the tuning never ran, and all 19 statistics of the one-year table sit inside the ranges real markets show. The 40 checks are graded over 21-year histories and the 19 statistics over one year, and in the preset table at the end 15/15 means 15 of the 15 fixed-roster statistics a record measured sit in band. How it is measured has every row. What it gets right Crash rates inside the registered band Over 90 simulated histories of 21 years, the index has 1.96 bear markets of 20% a decade against a real 1.12, 4.52 corrections of 10% against 3.65, and 9.6 sessions a decade down more than 5% against 6.2, all within the half-to-twice band the rows allow. Replaying 2008 and 2020 with the real VIX imposed, the maximum drawdown reads 45% and 37% against a real 57% and 34%. How often and how long the VIX stays high The VIX is above 30 on 5.9% of sessions against a real 8.2%, and a spell above 30 lasts 27 sessions on average against 22. Rates and bonds that move prices Driven through the real 2022 rate path, the market P/E moves -4.26% per 100 bp of the corporate yield, against the S&P 500's -5.2%. The simulated 10-year Treasury moves 4.96 bp a day against a real 5.41, and stocks and bonds co-move with the signs real markets show. Earnings that fall in recessions Around a contraction, aggregate earnings move -17.2% against -17% in Shiller's data, and after the packaged recession's low the index rises 55% within a year, against 69% after March 2009. Costs that grow with size The cost of an order grows with the square root of its size: the fitted exponent is 0.48, against the empirical 0.5. Leak checks A market that is easy to game makes a bad strategy look good and teaches a learning agent an edge real markets do not have, so some checks are strategies built to find leaks. Reading a headline 5 ticks late earns 15.8 bp, under the 20 bp limit. No price-only rule on the published suite of 20 markets beats buy-and-hold by more than 5 points, value and momentum screens stay inside bands set from real data, and timing the market on published economic data earns at most a point a year. pt-v19 fails both price-only checks. Long-run return and volatility The long-run index return is 6.4% a year against a target of 6.25%, with annual volatility of 19.1% against 18.1%. What a simulated market adds Fork a market at day 100, raise rates in one copy, and the difference between the copies is the effect of the rate rise on your strategy, with the luck held fixed. Run the same agent on 30 seeds and you have 30 versions of the same economy to test it on, each with crashes and recoveries of its own. A result says how a strategy behaves in this model. It does not forecast returns in a real market. What changed from pt-v19 The tape follows the model price. A company's own news moves its fair value for good. Plain market shocks are permanent up to a volatility ceiling, so the index no longer reverts on the mispricing's half-life. Fear marks fair value down while the VIX is above 40. Agents trade in a book with depth and a queue. The business cycle and GDP are published late, as the agencies publish them. Published at once, they would let an agent that went to cash on a contraction beat holding in every history. On the same pooled histories pt-v19 fails 16 of the 40 rows. It still runs, and replays exactly, with model="pt-v19". Limits The model still falls short in places. The named gaps lists each measured shortfall with the uses it rules out, and tf.envelope.check() refuses a question that depends on one. Every preset A shipped preset never changes, so a result names the preset it ran on and replays on it in every later release. Select one with model="pt-v19" or similar. The numbering skips 17. The counts are each record's fixed-roster panel of shape statistics, at one and two years. Preset Standing In band, one year In band, two years pt-v20 Default from 0.8.5 15/15 14/14 pt-v19 Default from 0.8.0, replaced in 0.8.5 15/15 14/14 pt-v18 Default from 0.7.0, replaced in 0.8.0 14/14 13/13 pt-v16 Default from 0.6.0, replaced in 0.7.0 14/14 13/13 pt-v15 Never the default 14/14 13/13 pt-v14 Default from 0.4.0, replaced in 0.6.0 14/14 13/13 pt-v13 Never the default 14/14 13/13 pt-v12 Default from 0.3.0, replaced in 0.4.0 14/14 12/13 pt-v11 Never the default 13/14 11/13 pt-v10 Default from 0.2.0, replaced in 0.3.0 13/14 12/13 pt-v9 Never the default 13/14 11/13 pt-v8 Never the default 13/14 12/13 pt-v7 Never the default 13/14 12/13 pt-v6 Never the default 12/14 9/13 pt-v5 Never the default 12/14 8/13 pt-v4 Never the default 11/14 8/13 pt-v3 Default from 0.1.0, replaced in 0.2.0 12/14 7/13 pt-v2 Never the default 11/14 8/13 pt-v1 Never the default 9/14 7/13 ====================================================================== # How it is measured https://docs.tradefloor.dev/how-its-measured.html How the default model is graded: checks registered before the run, exam seeds the tuning never saw, 40 rows over 21 years, the one-year table, the grade. ====================================================================== MODEL AND VALIDATION/ HOW IT IS MEASURED How it is measured Every check and its pass mark is written down and committed before the grade, the grade runs on seeds the tuning never used, and the scripts and their output are published. pt-v20 passed all 40 of its registered rows that way, and all 19 statistics of the one-year table are inside the ranges real markets show. Registration, then one grade 1 Measure real markets The S&P 500 since 1928, the VIX since 1990, 40 large US stocks, Treasury and corporate yields, Federal Reserve decisions and Shiller's long-run earnings give the figures a trader would notice, such as how many days a decade the index falls more than 5%. Each becomes a check with a pass mark wide enough to allow for how much the real figure moves from one decade to the next. 2 Write the rules Most behaviors come from a published source, such as the square-root cost of size. A new rule goes in switched off, and a test that runs one fixed market on every shipped preset proves that every other preset's output did not change by a single bit. 3 Screen on practice seeds Candidate settings run in grids over many simulated histories. Every market comes from a seed, and the seeds are split: practice seeds are used while choosing, and exam seeds are kept back for the grade. 4 Try to cheat it Some checks are strategies built to find leaks: trading on a headline 5 minutes late, rules that read only past prices, value and momentum screens, timing the market on published economic data, and trading on a rate decision. Each passes only if it earns about nothing, as it would in a real market. An independent audit found two leaks, and both were fixed in the model. 5 Register Every check, its pass mark and the exact settings to be graded are committed. Nothing on the list can change after the results come in. 6 Grade once The model runs once on the exam seeds and has to pass every row. If it fails one it does not ship, and a change needs a new registration and fresh seeds before another grade. 7 Freeze A model that passes ships as a preset with a fingerprint of its coefficients and known-answer digests that every release checks on five platforms. A shipped preset never changes. The seeds For pt-v20 the final settings were chosen on practice seeds 201 to 230, 501 to 530 and 801 to 830. The grade ran on exam seeds 101 to 130, 401 to 430 and 701 to 730: 90 histories of 21 years with nothing imposed, plus replays of 2008 and 2020 with the real VIX, the real 2020-21 and 2022 economies driven through the model, the packaged recession, and the leak-finding strategies. The 40 registered rows Each row is something a user would notice, with a tolerance that is easy to read: within 30% on the size of a crash, half to twice the real rate on how often something happens, and a hard limit on any edge a trading bot could learn. pt-v20 passes all 40. A rows replay 2008 and 2020, B rows count events and long-run figures over the free histories, C rows are leak and consistency checks and the cost of size, R rows cover rates, bonds and the driven 2022 market, S rows the packaged recession, D2, F1 and L1 the driven 2020 path, E1 earnings in a contraction, V1 variance over two and five years, and D1 the one-year table below. Where a row has both replays, 2008 comes before 2020. Row What it checks pt-v20 Real Pass mark A1 Worst month's volatility in the 2008 and 2020 replays within 30% of real 88.5 / 76.5 84.3 / 94.5 Within 30% A2 Maximum drawdown in the 2008 and 2020 replays within 30% of real 0.45 / 0.368 0.568 / 0.339 Within 30% A3 Peak stock correlation in the 2008 and 2020 replays within 0.15 of real 0.784 / 0.784 0.748 / 0.872 Within 0.15 B1 Share of sessions with the VIX above 30 0.059 0.082 1/2x to 2x B2 Mean length of a fear spell above VIX 30, sessions 27 22 1/2x to 2x B3 20% bear markets per decade 1.96 1.12 1/2x to 2x B4 10% corrections per decade 4.52 3.65 1/2x to 2x B5 Sessions down more than 5% per decade 9.6 6.2 1/2x to 2x B6 Share of sessions with the VIX under 15 0.4 0.326 1/2x to 2x B7 Index annual volatility, % 19.1 18.1 Within 20% B8 Long-run index return, % a year 6.4 6.25 Within 2 points of the target B9 Sd of annual index log returns, years 2-21, percent; the start-up drift: sd of the first 60 sessions' index return over the steady state's, across 50 markets 16.3 / 0.797 17.4 / 1 Annual within 20% of 17.4; start-up 2/3x to 1.5x C1 Crash rate in years 3-21 against years 1-2 0.92 1 2/3x to 1.5x C2 Histories touching the VIX ceiling, of 30 0 0 At most 1 C3 Edge from reading a headline 5 ticks late, bp 15.8 0 Under 20 bp C4a 65-minute lag-1 autocorrelation of print returns, median name; Roll spread over quoted −0.016 / 1.21 0 / 1 At or above −0.05, or Roll at most 2x quoted C4b The price-only rule furthest over its line on tf-suite-2026.1 (named in worst): median points over buy-and-hold, markets beaten of 20 0.2 / 11 0 / 10 Every rule at most +5 points and 14 of 20 C5 Stock-level (idiosyncratic) variance ratio at 60 sessions, median name, 30 histories of 2,660 sessions 0.95 0.924 In 0.80 to 1.05 C6 Rank IC of the value signal on published fundamentals against the next 20 sessions: whole history, first 60 sessions 0.0003 / 0.0048 0.0093 / 0.0093 Both in −0.03 to +0.05 C7 Rank IC of 12-1 and 6-1 month momentum against the next 20 sessions −0.0031 / −0.0013 0.0267 / 0.0413 12-1 in −0.04 to +0.095, 6-1 in −0.02 to +0.10 C8 Lo-MacKinlay one-day loser-minus-winner book, bp a day −0.0645 −1.74 In −6.4 to +2.9 C9 The cost of size in the agent-facing book: exponent and coefficient of the average cost against Q/V, in sigma units 0.484 / 0.424 0.5 / 0.5 Exponent in 0.4 to 0.7 and coefficient in 0.33 to 0.67 R1 Sd of the 2-year Treasury yield's daily change, bp (FRED DGS2 2015-2025) 3.87 5.23 In 3.65 to 6.80 R2 Sd of the 10-year Treasury yield's daily change, bp (FRED DGS10) 4.96 5.41 In 4.54 to 6.27 R3 Daily correlation of the roster index with a Treasury bond's return (minus the 10-year's change; SPY against IEF) −0.136 −0.161 In −0.36 to +0.03 R4 Daily correlation of the roster index with an IG bond's return (minus the corporate yield's change; SPY against LQD) 0.2 0.272 In +0.15 to +0.39 E1 The aggregate earnings fall around a contraction, median over the 30 histories' contractions (Shiller 1953-2020) −0.172 −0.17 In −0.40 to −0.046 D2 The driven 2020-21 market: the index's maximum drawdown and the sessions from the pre-crash high back to it, medians over histories 0.374 / 120 0.339 / 126 Drawdown in 0.237 to 0.441 and sessions in 63 to 252 F1 The driven 2020 path's fast crash: the high to the lowest close within 60 sessions, and the sessions it took 0.307 / 40 0.339 / 23 Drop in 0.237 to 0.441 within 12 to 46 sessions L1 Look-through: sessions from the index's trough to the aggregate earnings level's trough on the driven 2020 path 10.5 68 Lead in 1 to 136 sessions R5 The driven 2022 market: the index's maximum drawdown 0.265 0.254 Drawdown in 0.178 to 0.330 R6 The driven 2022 market: the market P/E's log change per 100 bp of the corporate yield, monthly averages −4.26 −5.2 −10.4 to −2.6 percent C10 Timing rules on published macro data against buy-and-hold, and the drift after a published turn 0.117 / 0.611 / 5 / −0.374 −2 / 0.324 / −5.9 / 4.3 Every rule at most +1.0 point a year and ahead in at most 2/3; drifts no larger than the S&P's after NBER turns R7a The equal-weight index's mean log move from the first price readable after a changed policy rate to the end of the first 65-minute bar, less the mean first bar of all days, bp: after a hike, after a cut −0.0853 / −1.73 0 / 0 Each within 5 bp, or 2 se where wider R7b The audit's rate-news agent through tf.evaluate: mean annual excess over holding in points, histories ahead −0.322 / 2 0 / 15 At most 0 points a year, ahead in at most 20 of 30 S1a The packaged recession: share of the index's paired log fall, at its lowest, won back 252 sessions later, mean over seeds 0.494 0.62 In 45% to 100% S1b The packaged recession: seeds whose cycle has left contraction and trough within 24 months of the onset 30 30 Every seed S2 The packaged recession: the index's own rise in the 252 sessions after its low, percent, mean over seeds 54.6 69 In +25% to +80% V1 The index's variance ratio at two and five years relative to one, years 2-21 of the pooled free histories 0.821 / 0.619 0.93 / 0.87 2y/1y in 0.75 to 1.15 and 5y/1y in 0.55 to 1.20 D1 Ruled bands of the one-year realism table in, on all four cells all all Every band in pt-v19, the previous default, fails 16 of the 40 rows on the same pooled histories. The rows are in tf.preset_record()["long_run"]. The one-year table The second test is a panel of statistics measured over one year, the median across 30 seeds, each against a band read from the longest real record the statistic allows. The 14 shape statistics and crisis dispersion are measured on a fixed roster of 40 companies; the four index rows on a roster that changes with each seed, because a crash rate measured on one roster describes only that roster. Statistic What it checks Real band pt-v20 In band annualised_vol_pct How much prices move in a year 12 to 41 20.5 yes excess_kurtosis How fat the tails are −13 to 24 18.1 yes return_acf1 Does yesterday predict today −0.07 to 0.06 0.013 yes abs_return_acf1 Does a wild day follow a wild day 0.02 to 0.17 0.0282 yes abs_return_acf5 The same, one week apart −0.03 to 0.1 0.0188 yes abs_return_acf20 The same, one month apart −0.05 to 0.06 0.0044 yes cross_sectional_corr How much names move together 0.09 to 0.49 0.305 yes volume_abs_return_corr Do big moves come with volume 0.35 to 0.64 0.596 yes leverage_effect Do falls raise volatility −0.11 to 0 −0.0341 yes volume_change_acf1 Does volume mean-revert −0.3 to −0.2 −0.268 yes corr_asymmetry Do names couple more when falling −0.15 to 0.23 0.0791 yes corr_asymmetry_lagged The same, one day later −0.15 to 0.33 0.086 yes sector_excess_corr Do industries move together 0.04 to 0.23 0.117 yes corr_persistence_acf1 Does correlation stay high after a panic −0.48 to 0.69 0.23 yes crisis_sector_dispersion Does one industry lead a crisis 0.79 to 1.74 1.3 yes index_drift_pct Which way the index goes 1.1 to 10.3 7.7 yes fear_gauge_dn1 How much fear a bad day buys 0.39 to 3.03 1.66 yes fear_gauge_dn3 The same, on a much worse day 2.6 to 9.58 4.26 yes index_tail_dn3_pct How often the index falls hard 0.64 to 2.34 0.89 yes The table certifies runs of up to 252 trading days. At 504 days pt-v20 holds 14 of 14 graded rows of the two-year panel, but tf.envelope.check() certifies one year, and the 40 rows above are the evidence for longer runs. Some statistics pass low: volatility clustering is weaker than real at every lag, and the named gaps below say what that rules out. Ask the package before you lean on a statistic: import tradefloor as tf verdict = tf.envelope.check(horizon_days=504, statistics=["abs_return_acf20"]) print(str(verdict).splitlines()[0]) OUTSIDE the envelope The verdict then names each reason: here the horizon is past the certified 252 days, and lag-20 volatility clustering is one of the named gaps. The named gaps Each gap says what the model gets wrong and which uses it rules out. tf.envelope.check() refuses a question that asks for one of them. Gap What falls short Do not use it for horizon The certified horizon is 252 days Multi-year backtests, and anything keyed on volatility dynamics beyond one year decay-shape Volatility memory is weaker than real at every lag Strategies whose edge depends on volatility clustering at any lag: volatility forecasts over one to five days, and vol targeting and risk parity on a one-month or longer estimate scenario-magnitude A driven scenario moves prices at a quarter to a half of the real size Sizing a scenario's impact rather than detecting it macro-range The endogenous macro state cannot reach its own crisis regimes Studying inflation regimes or policy crises from the endogenous economy alone roster-concentration A concentrated roster is measured on pt-v19 only, for four sector mixes and the shape rows Citing the certification for a concentrated roster on a level or crisis row, past 504 days, on any preset but pt-v19 (the default pt-v20 included), or for a sector mix other than the four measured These limits are measured too, and the envelope does not name them: Each session opens at the last print, so there are almost no overnight gaps and a stop held overnight is safer than it would be live. Nothing in the market learns your pattern and trades against you, and an order sliced over a day costs less than published studies of such orders find. With the VIX held at 65 the market is 5.1 times as volatile as at a VIX of 5, against a real 6.2. Some rows pass near an edge. The 2-year Treasury moves 3.87 bp a day against a real 5.23, near the floor of 3.65, and one macro timing rule uses 92% of its tolerance. The published grade The registration, the scripts that graded pt-v20, their inputs and the outputs of the grading run are in validation/pt-v20/ in the library repository, and validation/README.md says how to check the grade on a laptop or run it again. The one-year table is also published as data, in envelope.json, and docs/STATISTICS.md defines every statistic and gives the source of each band. What the grade does not say Passing means the model matches real markets on these figures and was not tuned to the exam. It does not mean a good score here predicts real returns. Some behaviors are not checked at all. Coefficients pt-v20 carries over from pt-v19 were chosen with a scoring rule over statistics that overlap the one-year table, so that table helped choose them and is weaker evidence than the rows graded on exam seeds. ====================================================================== # Release notes https://docs.tradefloor.dev/release-notes.html Every release and what it changed. An untraded run that names its preset replays exactly under later versions, and one that took the default does not. ====================================================================== RELEASES/ RELEASE NOTES Release notes A market with no agent orders in it replays exactly under every later release if the run names its preset. A run that took the default replays only under the version that shipped it, so each section below names the preset that moved. A traded run replays exactly on its own release, and one recorded before 0.8.5 matches only up to its first trade. The full history is in the library's CHANGELOG.md. 0.8.7 2026-10-03 A patch on the 0.8 long-term support line that changes text only. Every known-answer digest is 0.8.6's and pt-v20 stays the default, so every run replays exactly as it did. The documentation inside the package is rewritten for users. Each entry in Model parameters now says what the coefficient does, its units and which presets carry which value. The gap descriptions in the published envelope, the reasons score gives for a row it can't read, and several docstrings say the same in plain terms. A few stale claims were corrected on the way, and a test now checks every value a parameter's description states against the preset table. 0.8.6 2026-10-01 A patch on the 0.8 long-term support line, with no coefficient, default or trajectory changes. Every known-answer digest is 0.8.5's and pt-v20 stays the default, so every run replays exactly as it did under 0.8.5. tf.envelope.check() now says why it refuses a sector-concentrated roster on pt-v20. The four-mix roster run was repeated on pt-v20, on seeds 101 to 130 at 252 and 504 days. At 252 days every mix held every shape row the bands can grade. At 504 days the S&P-like and technology-heavy mixes put the correlation of volume with the size of a move at 0.6367 and 0.6332, against a ceiling of 0.63, where the balanced roster reads 0.6266. The refusal and the roster-concentration gap in How it is measured quote this run, which is in the library as measurements/roster-shapes-pt-v20.json. A question that passes preset="pt-v19" keeps the pt-v19 grant. 0.8.5 2026-10-01 Long-term support 0.8.5 starts the first long-term support line. The 0.8 line gets bug and security fixes for 24 months, and a fix ships in a 0.8.x patch only if it leaves every known-answer digest unchanged. The library's SUPPORT.md is the policy, and Support policy lists the digests. pt-v20 as the default A run that took the default will not replay against 0.8.1, so pass model="pt-v19" to keep that market. A run that names its preset replays exactly, and every shipped preset from pt-v1 on can still be selected. On pt-v20 the tape follows the model price, a stock's own news and the market's plain shocks move fair value for good, fear marks fair value down while the VIX is above 40, and agents trade in a book with depth. It passed all 40 long-run rows registered for it on exam seeds it had not been run on, all 19 statistics of the one-year table, and 14 of 14 at two years. How it is measured describes the grade. Nearest their edges are the two-year correlation of volume with the size of a move, at 0.627 against a ceiling of 0.63, a macro timing rule, at 92% of its tolerance, and the two-year yield, which moves 3.87 bp a day against a real 5.23. Economic data published late On pt-v20 the business-cycle phase reaches macro_fields["cycle"], trace rows and the LLM adapters' payload 252 sessions after it turns, as the NBER dates a turn about a year late. GDP growth is a quarterly figure released 21 sessions after the quarter ends. Prices read the true state, and code that needs it reads state_snapshot()["economy"], which a sandboxed agent cannot. No capture ratio on pt-v20 A shock on pt-v20 moves fair value for good, so knowing fair value leaves little edge and the Oracle's P&L follows the market's month. capture_ratio returns {} and warns why, versus_buy_and_hold gives each agent's P&L minus buy-and-hold's, and rank sorts on mean_excess_pnl. pt-v19 and earlier still report capture. Orders from act A number is a market order for that many shares. tf.Limit(quantity, price) and tf.Cancel() now work in evaluate and rank as they did in World. In evaluate a bad entry is refused with a line in the scorecard's errors and the rest of the dict trades. True and "100" as quantities, which traded before, are now refused, and a falsy return such as [] or 0 is an error where it passed as no trade. Return None or {}. Agents and evaluation has the rules. A read-only market for agents obs.engine is a MarketView and obs.portfolio a PortfolioView in evaluate, rank, World and tca.analyse. Anything else raises tf.SandboxError. An agent that changes the market or copies the engine is scored tampered, and rank leaves it out. privileged = True gives an agent obs.hidden, and trusted_agents=True hands every agent the live engine, and the scorecard marks both. The check catches writes and copies only. An agent can still read the seed and roster from the harness's frames, build a second engine and run it ahead, so run code you did not write in a separate process. Warm-up history evaluate, rank and World take history_days=N, which runs the market N days with nobody trading before the scored days. obs.history holds a daily bar per name and the published macro figures for every closed day, so a 20-day breakout rule can trade on day 1 instead of day 21. With history_days=0, the default, nothing changes. tca.analyse takes history_days too. Margin interest A negative cash balance pays the policy rate before each close in evaluate, rank and World, so a levered agent scores less than it did under 0.8.1. margin_interest=False borrows for free and the scorecard says free-borrowing. Explanation accuracy and its baseline The scorer asks which of ten factors moved prices most each day, reads the attribution after the close, and no longer counts fair_value_shift, which moves no price. On pt-v20 random_noise wins almost every day, so naming it every day scores 0.95 to 1.0. The scorecard carries explanation_baseline and explanation_edge and prints all three. Quote the edge. New scorecard fields The scorecard gains equity_curve, max_drawdown_pct, ruined, leverage_refusals, partial_fills, history_days, margin_interest and exposure_curve, the gross exposure after each step. Four read-only properties come from those curves: sharpe and volatility_pct, annualized from the daily returns with no risk-free rate subtracted, and time_in_market and avg_gross_exposure. The repr prints all four and counts errors. Agents and evaluation defines every field with its units. Refusals in rank An action an LLM adapter refuses on its own, such as a ticker the roster does not have, counts in Scorecard.rejected. AgentRecord.refused holds the count per seed, and the report gives the agent a REFUSED line with the first refusal. A seed on which an agent failed at every step has no score, so it no longer counts as ahead of a falling market. The report names the benchmark it read, and on pt-v20 it says that the Oracle ran and has no row. Bar volume Engine.bars() summed a running total, so a day bar read about two hundred times the day's volume. A bar's volume is now the shares traded in it, at every grain, so a day's tick rows add up to its five-minute bars and its day bar. If you read the tick column as a running total, take its cumulative sum per name and day. Prices and digests are unchanged. On pt-v20 volume_abs_return_corr moves from 0.508 to 0.596 at one year and from 0.561 to 0.627 at two, and volume_change_acf1 from -0.254 to -0.268 at one year, all still in their bands. LLM decisions and payload Decision schema 2 lets an action carry a limit_price, which becomes a tf.Limit, and side: "CANCEL", which becomes a tf.Cancel(). A bad action is refused on its own and the rest of the decision trades. Observation payload 1 is frozen for the 0.8.x line. portfolio.gross_exposure is renamed leverage, portfolio.open_orders lists waiting limit orders, and return_5d covers five full days. A recording carries both schema versions and a replay refuses one made under another, so recordings made before 0.8.5 do not replay. LLM adapters and MCP has the contract. Fingerprint battery version 2 Seven 120-day markets, one per shipped scenario including curve_shock, so a day-50 shock has 70 days after it. tf.battery(1) still builds the old six markets, and a fingerprint compares only with one taken on the same version. A digest for a traded run The known-answer script now hashes one run through tf.evaluate on pt-v20: the five reference agents and an agent that sends and cancels limit orders, with every order, fill and scorecard field. It is checked on all five platforms beside the other digests. Fills reach the market once Every harness passed an agent's fills to run_session as order_flow, which the session held on every tick, so one order counted 65 times and agents were marked to their own impact. They now arrive once, as fills. On pt-v19 the spec mean-reversion rule had beaten buy-and-hold on all 20 suite markets by a median 42 points in 60 days, and with fills applied once it reads +0.5 points, ahead in 10 of 20. Every traded result moves. run_session(order_flow=...) now raises, so pass fills=, or flow_per_tick= for a standing rate. Logs, checkpoints and manifests from 0.8.x replay as they ran. Engine and data has the change. A book with depth Seven ModelParams dials give the book depth priced by size, consumption and refill, resting limit orders and a permanent impact per agent. All seven are 0 on every earlier preset, and pt-v20 turns them on. Engine and data says what each does to an order. A resting order fills against the model's flow only at a price inside the maker's quote for that tick. A resting order that the maker's re-quote crosses trades at the maker's price and is recorded as liquidity="taker". tf.Limit and tf.Cancel compare by value, so agree no longer reports two forks that sent the same limit orders as different. Rate indices Universe.random(40, seed=1, bonds=True) adds UST2Y, UST10Y and IGCORP, priced off the engine's own curve. None is a real security, and every equity price is the same with or without them. The curve_shock scenario moves the whole curve 200 bp in one day, baselines.Balanced holds a 60/40 book, and cash_interest=True pays idle cash the policy rate. Engine and data has the pricing. 64-bit seeds Seeds take any 64-bit integer, and every seed below 2**32 gives the market it gave before. Draw a sealed seed with secrets.randbits(64). Recalibrated crisis scenarios On pt-v20 recession takes the index down 44.7% at 120 sessions (2008 fell 45%) and now ends, with the cycle at trough on day 365. liquidity_crisis falls 33.9% at worst, as March 2020 did, though more slowly. Scenarios says why. Gym resets The gym environment draws a new market on each reset(), where a reset without a seed used to replay the constructor's market. info["seed"] names the episode's seed. An action over the leverage cap is scaled down to 1.96x gross, with info["scaled"] set. Speed evaluate of five momentum strategies on 40 names over 20 days fell from 11.0 to 3.95 seconds of CPU, and building a pt-v20 engine from 0.70 to about 0.01 seconds. Python 3.12 and 3.13 3.12 changed sum() over floats, so an agent's orders could split from 3.11's in the last digit. tradefloor now adds floats left to right on every version, so 3.12 and 3.13 match 3.11. Short-lag clustering Volatility clustering reads about a quarter of real at lag 1, so the decay-shape gap now covers abs_return_acf1 and abs_return_acf5 as well as abs_return_acf20, and tf.envelope.check() refuses a question that relies on any of the three. How it is measured lists the gaps. Examples and the MCP server examples/08-claude-agent.py replays a committed Claude run by default, so it runs with no key and no network, and TRADEFLOOR_LIVE_EXAMPLES=1 with a key calls Claude once per simulated day. The MCP server refuses a stress run that ends before its scenario's first event, where it returned a difference of 0.0, and refuses strategies named after a baseline. list_scenarios gives each scenario's first and last event day, the constructors are rate_ramp and vix_shock, and pip install "tradefloor[mcp]" now installs pyarrow, which explain needs. The model written down The model specification states pt-v20 as equations, each with its source line, and gives every coefficient its value and how it was set. STATISTICS.md names the statistic sets behind every count the site quotes. Rust crate API changes The crate takes the package's version, so Cargo treats 0.8.5 as a compatible update to 0.8.1, and it is not. Seeds are u64, several public structs have new fields and are #[non_exhaustive], and Engine::new builds a pt-v20 market. Pin tradefloor = "=0.8.1" to stay on the old API. The CHANGELOG lists every changed signature. Smaller fixes An agent no longer trades with itself in the settlement book. A resting fill can sit outside the day's high and low, because a bar keeps only each tick's last print. The Scorecard repr prints sharpe=n/a (short run) for a run of fewer than 20 scored days. An act() that returns something other than a mapping is reported on an UNUSABLE line and held in AgentRecord.unusable. The scripts and output that graded pt-v20 are in the library's validation/pt-v20/. Known gaps The worst month of the 2020 replay is about 19% milder than the real one. With the VIX held at 65 the market is 5.1 times as volatile as with it held at 5, against 6.2 times in real markets. 0.8.1 2026-09-23 Text only, with no coefficient, default or trajectory changes. The README, the messages tf.envelope.check() prints, the scenario target notes and the parameter descriptions now describe pt-v19, where some still quoted older presets. Re-measured on pt-v19, the decay slope of volatility clustering reads -0.515 against a real -0.436, and the memory holds to lag 20. A driven 2020-21 scenario moves prices at about a fifth of the real size, and qe_pe_boost moves nothing on pt-v16 and later. 0.8.0 2026-09-23 pt-v19 became the default. A run that took the default will not replay against 0.7.x, so pass model="pt-v18" to keep the 0.7.x market. pt-v19 meets all fifteen criteria of a long-run check (thirty 21-year histories, and 2008 and 2020 replayed with the real VIX forced in), where pt-v18 meets eight. tf.preset_record("pt-v19")["long_run"] holds every row. Over 21 years it has 1.35 bear markets a decade against a real 1.12, index volatility of 16.6% against 18.1%, and the VIX above 30 on 8.1% of sessions against 8.2%. The 2008 replay falls 41% against the real 57%. ModelParams.from_preset refuses an override that breaks one of pt-v19's three identities, such as garch_alpha=0.07 alone, because garch_beta follows from it. ModelParams.from_preset_unchecked builds it anyway, and ModelParams.identity_breaks lists what broke. A recorded transcript names its preset under meta["model_preset"], and replaying it against a different preset raises ReplayMiss. A manifest or checkpoint written under 0.7.x is refused, because its probe simulation runs the default preset. state_hash covers thirteen more fields, and a custom parameter set's fingerprint changes between releases. Realism is graded on the ruled bands, each read from the longest real record for its row. envelope.certified(), envelope.score() and facts.report() take basis="shipped" for the old table. New calls: Engine.session_news(), Engine.session_tick, Scenario(vix_sets_variance=True) and hold(epicentre=...). 0.7.1 2026-09-08 Documentation only. No behavior changes, and every published digest holds. 0.7.0 2026-09-08 pt-v18 became the default, and the first default to hold every certified row: the index returns +5.80 percent a year where pt-v16 lost 13.64. The steady-state crisis lever reads 6.53x against real markets' 6.16x. tf.preset_names() lists the presets the engine resolves, and tf.preset_record reads pt-v1 through pt-v18 from JSON in the wheel. Every Universe.random roster re-rolls, so pin 0.6.2 for the old draws. An EDGAR roster and a roster you build yourself stay as they were. A supplied macro_state now survives construction, the first central-bank meeting and OPEC decision fall inside a 252-day year, and to_instruments reads the discount rate from the model. Engine.explain decomposes a name's day down to the draws that seeded it. 0.6.2 2026-09-01 A scenario's liquidity shock now reaches the agent's volume and order cap. World(on_refusal="skip") records an unusable response and carries on. Every adapter takes prior=, a recording consulted before the provider. resample() measures an agent's own noise floor. fetch(ciks=) returns exactly those EDGAR filers, and saved files are the same bytes on Windows. 0.6.1 2026-08-31 tradefloor.integrations gains adapters for the OpenAI Agents SDK, PydanticAI and LangGraph, and one for any plain Python function. A Transcript keys each exchange by a digest of the exact input, so a recorded run replays with no framework installed and no network. The examples in examples/integrations/ run offline. LLM adapters and MCP has the contract. 0.6.0 2026-08-30 pt-v16 became the default. It couples the VIX to the market's own realized volatility. tf.Scenario and six packaged scenarios arrived, with the tradefloor scenario command. tradefloor.counterfactual runs one agent in two worlds that differ by one variable, with World, agree() and compare(). 0.5.0 2026-08-28 The library was renamed from pretium, which published through 0.4.3 and stays on PyPI and crates.io. A result computed under one of those versions is cited as pretium at that exact version. The rename changes no behavior, and preset names stay pt-v1 through pt-v15. pt-v15 is selectable by name, and pt-v14 remains the default. 0.4.3 2026-08-28 Restores pt-v13 and pt-v14 exactly as they were in 0.4.0 and 0.4.1, after 0.4.2 moved their dollar safe-haven gate. The gate is its own dial now, usd_crisis_vix_threshold. 0.4.2 2026-08-28 Three fixes with no trajectory change: crisis_vix_threshold now gates the dollar's safe-haven drift as well as gold, DayAdvanceOutcome carries a meeting's decision, and daily_credit_floor_gain ships at 0.0. 0.4.1 2026-08-28 pt-v13 and pt-v14 now report the 60-day mispricing half-life they run, where tf.model_preset() said 68.26. No trajectory moves. 0.4.0 2026-08-28 pt-v14 became the default. Over 13 seed blocks it holds the two-year panel fully in band on 11, where pt-v12 held 3, and puts 137 of 138 roster shapes in band. 0.3.0 2026-08-27 pt-v12 became the default, the first preset in band on all 14 statistics over two years as well as one. Results recorded without naming a preset differ from here on, and model="pt-v10" gives the old market. ====================================================================== # Support policy https://docs.tradefloor.dev/support.html Which tradefloor releases get fixes and for how long, what a patch release may change, how presets are frozen, and the digests each release checks. ====================================================================== RELEASES/ SUPPORT POLICY Support policy tradefloor is before 1.0, so the API can change in any minor release. Pin the version you tested with: pip install "tradefloor==0.8.7" The model is versioned separately, as presets, so a result names both: the package version and the preset it ran on. This page is a summary of SUPPORT.md in the library repository, which is the policy itself. Supported releases Version Supported 0.8.5 and the 0.8 patches after it Yes. The first long-term support line, with pt-v20 as the default. 0.8.1 and earlier No fixes. Every release from 0.5.0 on stays installable from PyPI and crates.io. The 0.8 line gets bug and security fixes for 24 months from the day 0.8.5 was tagged. After that it stays installable and gets no more fixes. The next long-term support line is named at least 6 months before this one ends, so the two overlap. Old releases are never removed, and a release is never yanked to hide a model change, because published results replay on the version that made them. To report a vulnerability, use the private security advisory form on GitHub. Changes allowed in a patch release A 0.8.x patch may ship: a bug fix that leaves every known-answer digest unchanged a security fix, under the same condition wheels for a new CPython version or platform, built from the same source and giving the same digests documentation and error messages These stay fixed for the life of the 0.8 line: every shipped preset's coefficients, and so its fingerprint the known-answer digests, because a digest that moves means the market moved the default preset, pt-v20 the random draw schedule: how many draws are taken, from which stream and in what order the public API of the Python package and the Rust crate. Nothing is removed or renamed, no signature changes in a way that breaks a call, and no new features are added saved formats. A checkpoint, RunManifest or recorded transcript written by one 0.8 patch loads in every other what an LLM agent is shown and how it answers: observation payload version 1 and decision contract version 2 A defect that would change a market cannot ship in a patch. It is written up as an erratum in the release notes and in tradefloor.envelope, where tf.envelope.check() can refuse the questions it affects, and the fix ships in the next minor release, as a new preset if it changes coefficients. Presets across releases A preset is frozen when it first ships in a tagged release. From then on its coefficients, its fingerprint and its known-answer digest stay the same, and a later release can measure it again and publish new figures about it without changing what it computes. A better coefficient becomes a new preset with a new name. From the 0.8 line on, a new preset becomes the default only in a minor release such as 0.9.0, never in a patch. 0.8.5 was the one exception, because the policy starts there. The release notes say which preset moved and what changes for a user. A preset that is no longer measured or recommended is retired. It keeps running exactly and stays in the package, because removing it would break every result that cites it. Reproducibility says what a preset name pins and what it does not. Release checks Every release builds wheels for five targets (Linux x86_64 and aarch64, macOS arm64 and x86_64, Windows x86_64), runs the known-answer tests on each and stops if any target's digests differ from the others or from the committed ones. The tests hash these runs, one digest each: sim one fixed untraded market of twelve companies over five sessions on the default preset meta the coefficients the library reports for that preset, kept apart so a reporting fix cannot pass for a market change bonds the same market on a roster with the three rate indices presets one 60-session untraded market on every shipped preset, each hashed on its own highseed a market on a seed above 2**32 traded the reference agents and an agent that sends and cancels limit orders, through tf.evaluate on pt-v20, hashing every order, fill and scorecard field book the order book on a fixed script of market and limit orders and cancels For 0.8.5 and 0.8.6 the simulation digest begins 72485a9f, the preset digest 87f0b185, the traded digest 8e032d38 and the book digest b14d1f50. The presets and traded digests are new in 0.8.5. The release notes list what changed in each version. ====================================================================== # Citing tradefloor https://docs.tradefloor.dev/cite.html How to cite tradefloor in a paper: the BibTeX entry, the version and preset to name, and what to publish beside a result so a reader can rebuild the market. ====================================================================== RELEASES/ CITING TRADEFLOOR Citing tradefloor Cite the version you ran and name the preset. One version can run every shipped preset, and results depend on which one you used. @software{tradefloor, author = {Coombes, Simon}, title = {tradefloor: a deterministic market simulator with a limit order book}, version = {0.8.7}, year = {2026}, url = {https://github.com/simoncoombes/tradefloor}, note = {Model preset pt-v20} } There is no DOI yet, so the version and the repository address identify the software. The repository's CITATION.cff carries the same details, and GitHub's "Cite this repository" button reads it. In the text, name the model as well, for example: "tradefloor 0.8.7, preset pt-v20, specified in its docs/MODEL.md". How prices are made describes that model. Details to publish with a result A citation identifies the software. To let a reader rebuild your market, report these too: the tradefloor version and the preset name the seeds the universe: the Universe.random(n, seed=...) call or the roster file, and its fingerprint the economy on day zero, if you set one any scenario file, and its digest A RunManifest records all of these and replays the market. It checks the market and carries no score, so to let a reader check a score, publish the agent's code, the tf.evaluate or tf.rank call and the seeds as well. Reproducibility says what each of these inputs fixes. License tradefloor is open source under the MIT license or the Apache License 2.0, at your option.