Skip to the page
GETTING STARTED/TROUBLESHOOTING

Troubleshooting

Error text is quoted as tradefloor 0.8.7 prints it, and numbers in it vary from run to run.

Installation and extras

No matching distribution

Symptompip install tradefloor ends with No matching distribution found for tradefloor.
CauseThe Python running pip is older than 3.11, the oldest version tradefloor supports.
FixCheck with python --version, then install into a 3.11 or later environment, for example python3.12 -m venv .venv.

A build from source

Symptompip downloads a .tar.gz instead of a wheel, and the build fails with an error about Rust, cargo or maturin.
CauseThere is no prebuilt wheel for your platform. Wheels cover Linux x86_64 and aarch64, macOS arm64 and x86_64, and Windows x86_64, and anything else builds from source.
FixInstall a Rust toolchain from rustup.rs and run pip again, or use one of the five platforms (Install).

Brackets in zsh

Symptompip install tradefloor[mcp] stops with zsh: no matches found: tradefloor[mcp].
Causezsh reads the square brackets as a file pattern before pip sees them.
FixQuote the argument: pip install "tradefloor[mcp]".

Missing extras

SymptomAn import or a call raises an ImportError that ends in a pip install command, such as tradefloor.gym needs numpy and gymnasium, and numpy is not installed. Install them with: pip install 'tradefloor[rl]'.
CauseThe feature needs an optional package that the core install leaves out. The MCP server, the gym environment and the live mode of each LLM adapter work this way.
FixRun the command in the message. Optional extras lists every extra.

FinRobot on Python 3.12 and later

Symptompip install "tradefloor[finrobot]" stops with Could not find a version that satisfies the requirement finrobot>=0.1.5.
CauseFinRobot supports Python 3.10 and 3.11, and tradefloor needs 3.11 or later, so the extra installs on 3.11 only.
FixUse a Python 3.11 environment for live FinRobot runs. Replaying a recorded FinRobot run needs no extra, on any supported Python.

An agent that makes no trades

The scorecard reads trades=0 and a P&L of 0. Look at the end of the scorecard first: errors=N there means the agent tried and something failed.

Acting on the wrong step

SymptomNo errors, no trades, or trades only on the first day.
Causeobs.step counts decision steps over the whole run, from 0 to six times the number of days, less one. A check written as obs.step == 0 to mean "every morning" is true once. obs.step_of_day starts again at 0 each day.
FixUse obs.is_first_step_of_day for once a day and obs.step == 0 for once a run.

Errors on every step

SymptomA warning such as Agent 'mine' failed on all 120 of its steps, so its score is empty. First error: ..., and errors=120 on the scorecard.
Causeact raised, or returned something other than a mapping. A list of pairs gives act() must return a mapping of ticker to order, such as {'AAA': 100} or {'AAA': tf.Limit(100, 25.0)}, or None to trade nothing. It returned a list: ....
FixPrint scores["mine"].errors, where each line names the step and what went wrong. A step that raised traded nothing, and the steps after it ran as normal.

A class passed as the agent

Symptomtf.evaluate raises before any market runs: Agent 'mine' is the class Mine, not an agent. Pass Mine() instead.
CauseThe agents mapping holds the class, and the harness needs an object with an act(obs) method.
FixPass an instance: {"mine": Mine()}.

Limit orders that never fill

SymptomNo errors, and trades stays at 0 while the agent sends tf.Limit orders.
CauseA limit order waits in the book until the market reaches its price, and trades counts fills, so an order that never fills adds nothing. A buy limit well below the current price can wait for the whole run. A new limit on the same ticker replaces the one waiting there.
FixSet the limit price from obs.price(ticker), or send a plain number of shares for a market order that fills at once.

An agent that needs past prices before it acts can also sit out a whole run, which Missing warm-up history covers.

Rejected orders

A refused order trades nothing, adds one to the scorecard's rejected and adds a line to errors, and the other orders from the same step still trade.

Message in errorsCauseFix
step 4: no instrument with ticker "aaa" in this universeThe ticker is not in the market. Tickers are case-sensitive.Take tickers from obs.tickers.
step 2: trade would take leverage to 2.37x, above the 2.00x limitThe order would take gross positions past max_leverage times net worth, 2.0 by default. Counted again in leverage_refusals.Send a smaller order, or pass max_leverage to tf.evaluate. None removes the limit.
step 0: the order for 'AAA' must be a number of shares, a tf.Limit or a tf.Cancel, got '100' (str)The quantity is a string or a bool.Send an int or a float.
step 0: the order for 'AAB' must be finite, got nanA calculation produced NaN or infinity, often a division by zero.Check the inputs to the size calculation, and send nothing when they are missing.

Once an agent's net worth reaches zero or below, the leverage limit refuses every order, because leverage over a net worth of zero is infinite.

A market order larger than the book can absorb is not refused. It fills what the book holds, and the scorecard lists it in partial_fills, as in step 2: asked to buy 1,000,000,000,000 AAD; the book held 44,021,201, and the rest did not fill.

Shares and portfolio weights

SymptomThe agent trades but the P&L is a few units of currency, with a warning such as Agent 'mine' asked for fractions of a share (0.2 of AAA, 0.2 of AAB, 0.2 of AAC, ...). act() returns numbers of shares, not portfolio weights.
Causeact returned portfolio weights. {"AAA": 0.2} buys a fifth of one share of AAA.
FixConvert each weight to shares at the current price, as below.

Each value act returns is a trade to make now, in shares, and not a position to hold. Returning {"AAA": 100} at every step buys 100 more shares at every step. To hold a target weight, trade the difference between the target and the current position:

def act(self, obs):
    worth = obs.portfolio.net_worth()
    orders = {}
    for ticker, price in zip(obs.tickers, obs.prices):
        target = int(0.2 * worth / price)             # 20% of net worth
        orders[ticker] = target - obs.position(ticker)
    return orders

tf.baselines.rebalance(obs, {"AAA": 0.2, "AAB": 0.2}) does the same, skips trades of less than one share and caps each trade at 2% of the company's average daily volume, which limits what the agent pays in impact.

Missing warm-up history

SymptomAn agent that needs, say, 20 days of prices trades nothing for the first 20 days, or nothing at all in a short run. obs.history.bars("AAA", last=20) returns fewer than 20 bars, and none at the first step.
Causetf.evaluate and tf.rank start at day 0 with no history unless asked for one.
FixPass history_days=20. The market then runs 20 untraded days before day 0, and obs.history holds their daily bars at the first decision, labeled day -20 to day -1.

A warm-up changes which days are scored. The scored days continue the warmed market, so on the same seed they are different days from a run without one, and the scorecard records history_days to say so. A warm-up can be up to 2520 days, ten 252-day years. The day in progress is never in obs.history, and each scored day joins it after its close. The reference agents and tf.StrategySpec strategies keep their own price history and ignore obs.history.

While the history is still empty, obs.history.bars() returns an empty list even for a ticker the market does not list, so check tickers against obs.tickers.

Replay mismatches

Manifest refusals

RunManifest.reproduce() raises a ValidationError naming the part that disagreed.

Message starts withCauseFix
this build does not reproduce the manifest's eraThe manifest was written on a release whose default preset or arithmetic differs from this one. The message names the version and platform that wrote it.Install that version, or rebuild the run from its named preset, seed and universe.
this manifest ran model preset 'pt-vN', which this build does not shipThe preset is newer than this release, or came from a development build.Install a release that ships the preset.
model preset 'pt-vN' disagrees between this manifest and this buildSame name, different values: one side was a development build (Preset names and preset values).Reproduce on the build that wrote the manifest.
the replay ran but did not rebuild the recorded marketThe inputs matched and the market did not. If the message says the draw counts differ, the order log does not cover everything that happened to the engine.Replay part of the log with tf.replay(log, ..., until=n) and find the first step that differs.
the universe in this manifest does not match its recorded fingerprintThe file was edited after it was written.Get an unedited copy.

A traded run recorded before 0.8.5 replays up to its first trade and differs after it on 0.8.5 and later, because 0.8.5 changed how fills reach the market. Replay it on the release that recorded it.

Refused LLM replays

A replay looks up each answer by a digest of the exact observation the model was shown. When it cannot find one it raises ReplayMiss, which stops tf.evaluate with a note naming the step, the agent and the seed, so it is never scored as an agent that held cash.

Message starts withCauseFix
no recorded response for step N (day D, digest ...)Something that feeds the observation changed after recording: the seed, the universe, history_days, the instructions or the market settings. A recording made before 0.8.5 can fail this way too, because 0.8.5 changed the observation the model is shown.Replay with exactly the call that recorded it, on the release that recorded it, or record again live.
this transcript was recorded against a different simulation presetThe replay runs another preset, often the default.Pass the recorded preset, as in tf.evaluate(..., model="pt-v19").
this transcript was recorded under observation payload versionThe recording was made by a release with another observation payload. Since 0.8.5 each recording names its payload version.Replay on the release named in the message.
the recorded entry for step N (day D, digest ...) holds a null responseThe live call failed during recording, and the failure was recorded.Record the run again live.

Model-provider failures

Failed model calls

SymptomThe scorecard has errors=N and few or no trades, with lines such as step 12: FrameworkError: openai-agents raised APIConnectionError instead of returning a decision: ....
CauseThe call to the model never completed: no network, a missing or wrong API key, a rate limit or a timeout. The adapters add no retry of their own, and tf.evaluate scores the step as one that traded nothing and carries on.
FixRead the exception type at the end of the line, fix the key, quota or connection, and run again. Compare runs by their errors count as well as their P&L, because a failed call and a decision to hold both trade nothing.

Unreadable decisions

SymptomLines in errors such as DecisionError: no JSON object in the framework response, no 'actions' key in the decision (keys present: ...) or the framework returned an empty response, so there is no decision to validate.
CauseThe model answered, and the answer was not a decision: an actions list and an optional rationale. The adapter refuses to guess what was meant.
FixAsk for that shape in the instructions, or use the framework's structured output. An answer of {"actions": []} is a valid decision to hold.

Refused actions

SymptomThe decision traded, and errors and rejected also count one action, for example a symbol the market does not list.
CauseSince decision schema 2 a bad action inside a good decision is refused on its own, and the rest of the decision trades.
FixRead the refused action: line in errors, which gives the reason.

The PydanticAI request budget

SymptomUsageLimitReached: the run at step N (day D) stopped on its request budget of ...
CauseThe PydanticAI adapter caps the requests one decision may make, at 8 by default.
FixRaise request_limit on the adapter, or pass None to remove the cap.

Adapter warnings

MessageCauseFix
... is starting a new run (day 0, step 0) but still holds N steps of prices from an earlier oneOne adapter ran a second evaluate or World, so its model sees the first market's prices in its returns.Build a new adapter for each run, as tf.rank does with its factory.
this agent has output validators registered with @agent.output_validatorPydanticAI refuses to change the output type of an agent with an output validator.Pass bind_output_type=False to keep your output type. tradefloor still validates the decision.
GraphInterruptedErrorA LangGraph graph paused for a human. The market moves on as soon as act returns, so there is nothing to resume.Give the adapter a graph that decides without an interrupt.