Skip to the page
GETTING STARTED/REPRODUCIBILITY

Reproducibility

A tradefloor market is a function of its inputs. Run it twice with the same inputs on the same release and you get the same market to the last bit, on any of the five platforms tradefloor ships wheels for.

The inputs

These fix a market. Change any one of them and you get a different market.

  • The universe: the companies, in order. The engine draws random numbers in roster order, so the same companies in another order are a different market. universe.fingerprint, a sha256 of the roster, covers the order as well as the companies.
  • The economy on day zero: the macro_state passed to tf.Engine. Leaving it out is an input too. The engine then settles its own default opening, which is a different market from passing tf.Macro() with the same values.
  • The seed: the seed passed to tf.Engine, any integer from 0 to 2**64 - 1. Universe.random takes a separate seed, which picks the companies and not the market.
  • The preset: the model passed to tf.Engine or tf.evaluate. Leaving it out uses the installed version's default, which is pt-v20 in 0.8.7.
  • The package version, which decides what the default preset is and which code runs.
  • Any scenario applied to the market, day by day.
  • For a traded run, the agent's orders, in the order they were sent, and for an LLM agent the answers the model gave.

The snippets on this page run in order in one Python session. This one runs the same 20-day market with one input changed at a time and compares a digest of the whole engine state:

import tradefloor as tf

universe = tf.Universe.random(5, seed=11)


def final_state(roster=universe, seed=42, macro=None, model=None):
    """Run 20 days and return a digest of the whole engine state."""
    engine = tf.Engine(seed=seed, universe=roster, macro_state=macro, model=model)
    engine.run_days(20)
    return engine.state_hash()


reference = final_state()
print("same inputs again:  ", final_state() == reference)
print("roster reversed:    ", final_state(roster=tf.Universe(reversed(universe))) == reference)
print("VIX 30 on day zero: ", final_state(macro=tf.Macro(vix=30.0)) == reference)
print("seed 43:            ", final_state(seed=43) == reference)
print("preset pt-v19:      ", final_state(model="pt-v19") == reference)
same inputs again:   True
roster reversed:     False
VIX 30 on day zero:  False
seed 43:             False
preset pt-v19:       False

Markets with and without an agent

A market with no agent orders in it depends on the inputs above and nothing else. The same inputs give the same market, bit for bit, on Linux, macOS and Windows. tradefloor carries its own exp, log, pow, sin and cos, so the platform's math library cannot change a result, and every release runs fixed simulations on all five wheel targets and stops if any digest differs.

A run with an agent in it has one more input. The agent's orders fill against the order book and move prices, and the moved prices are what the agent sees at its next decision. So the agent's decisions are part of the input, and two agents on the same seed trade two different markets after their first trade.

  • A Python agent that decides the same way from the same observations gives the same run again on the same release. Anything in it that varies, such as an unseeded random number or a value read from the clock, changes its orders and so the market. Its own arithmetic is outside tradefloor's checks too. Python 3.12 changed how sum() adds floats, so an agent that sums prices with sum() can send different orders on 3.11 and 3.12.
  • An LLM agent can answer differently every time it is asked. The LLM adapters record every call and its answer in a transcript, and a replay feeds those answers back without calling the model. A replay looks each answer up by a digest of the exact observation the model was shown, so it stops with ReplayMiss when anything upstream has changed. Replay mismatches lists the messages.

Every input an engine consumes, orders included, is in engine.order_log, and tf.replay(log, seed=..., universe=...) rebuilds the market from that log without the agent's code.

Across releases

A shipped preset never changes. Its coefficients are fixed when it first ships in a tagged release, and a better coefficient becomes a new preset with a new name. So an untraded market on a named preset replays exactly in every later release, and every shipped preset from pt-v1 on can still be selected with model="pt-v12" or similar. Each release from 0.8.5 on checks one known-answer digest per shipped preset on all five targets.

The default preset moves between releases. pt-v20 became the default in 0.8.5, replacing pt-v19, so a run that left model out on 0.8.1 gives a different market on 0.8.5 and later. Name the preset in anything you publish.

Traded runs have a narrower promise. In 0.8.5 and later an agent's fills reach the market once, on the next tick, where 0.8.1 fed them in as order flow on every tick of the next step. That change applies to every preset, so a traded run recorded before 0.8.5 matches up to its first trade and differs after it. Each release also checks one traded run on all five targets: the reference agents through tf.evaluate on pt-v20, with their orders, fills and scorecards. Runs with your own agents go through the same code. Within the 0.8 long-term support line a patch release has to leave every one of these digests unchanged (Support policy).

Preset names and preset values

A preset name is fixed only in tagged releases. A development build or an edited source tree can carry other values under the same name. A RunManifest and a transcript both store the values and refuse a mismatch, and a coefficient changed through the API, as in tf.ModelParams.from_preset("pt-v20", momentum_theta=0.05), fingerprints as custom- and eight hex digits, never as the preset. To pin a model, install a tagged release from PyPI or crates.io and name the preset.

Saving a run

A RunManifest holds what a stranger needs to rebuild a market: the version, the preset and its coefficients, the seed, the universe, the economy on day zero, any scenario and the order log, with a digest of the market the run ended on.

engine = tf.Engine(seed=42, universe=universe)
engine.run_days(20)

text = tf.RunManifest.of(engine, seed=42, universe=universe).to_json()
rebuilt = tf.RunManifest.from_json(text).reproduce()   # replays, then checks
print(rebuilt.state_hash() == engine.state_hash())
True

reproduce() checks every component against its fingerprint before it replays, and raises an error naming the part that disagrees. It first runs a small fixed probe simulation on the default preset, so a manifest written on a release with a different default (0.8.1, read on 0.8.5 or later) is refused. Rebuild such a run from its named preset, seed and universe instead.

A manifest checks the market and carries no score. tf.evaluate and tf.rank write no manifest, so to let a reader check a score, publish the agent's code, the call that scored it and the seeds. Citing tradefloor lists what to report beside a result.