Write an agent
An agent is any Python object with an act(obs) method. This guide builds one, sizes its orders, reads what it is allowed to see and gives it price history before its first decision. Every argument, default and return value is in the reference, Agents and evaluation.
Every code block on this page runs on its own.
A first agent
At each decision step, six a day by default, the harness builds an observation, passes it to act and reads back a dict of ticker to order. tf.evaluate runs the agent on a market built from a seed and a universe, here 40 generated companies, and returns a scorecard per agent.
import tradefloor as tf
market = tf.Universe.random(40, seed=111)
class Quoter:
# bid a cent under the last price at each open, cancel what is left at the close
def act(self, obs):
if obs.is_first_step_of_day:
bid = round(obs.price("AAA") - 0.01, 2)
return {"AAA": tf.Limit(500, bid)}
if obs.is_last_step_of_day:
return {"AAA": tf.Cancel()}
return None
card = tf.evaluate({"quoter": Quoter()}, seed=7, universe=market, days=10)["quoter"]
print(card)Scorecard('quoter', pnl=20,597, return=+2.06%, trades=39, impact=+1.47bps, sharpe=n/a (short run), vol=18.4%, in_market=100%, exposure=0.56x)trades counts fills. Ten limit orders made 39 of them, because a waiting order fills in parts as the price comes to it. Each agent starts with 1,000,000 in cash and may hold positions worth up to twice its net worth.
obs.step counts decision steps across the whole run, so day 1 starts at step 6. A guard written if obs.step == 0 therefore fires once a run, and an agent that means once a day reads obs.is_first_step_of_day or obs.step_of_day instead.
Orders from act
act returns a dict of ticker to order, or None to trade nothing that step. A plain number is a market order for that many shares, positive to buy and negative to sell or short. tf.Limit(quantity, price) waits in the book at its price or better, and tf.Cancel() withdraws every waiting order on the ticker. A new tf.Limit on a ticker replaces the one waiting there. The full table, with what happens to a partial fill, is in What act returns.
A bad entry is refused on its own and the rest of the dict still trades. Each refusal adds a line to the scorecard's errors, so read errors before the P&L:
import tradefloor as tf
market = tf.Universe.random(40, seed=111)
class Careless:
# a quantity sent as text, and a ticker the market does not list
def act(self, obs):
if obs.step == 0:
return {"AAA": 100, "AAB": "100", "ZZZZ": 50}
return None
card = tf.evaluate({"careless": Careless()}, seed=7, universe=market, days=1)["careless"]
print(card.trades, card.rejected)
for line in card.errors:
print(line)1 2 step 0: the order for 'AAB' must be a number of shares, a tf.Limit or a tf.Cancel, got '100' (str) step 0: no instrument with ticker "ZZZZ" in this universe
If an agent's orders never fill or are refused, An agent that makes no trades and Rejected orders list the usual causes.
Sizing to a target weight
Values in the dict are shares, never weights. To hold 20% of net worth in AAA, the agent works out the holding it wants, 0.2 × net worth ÷ price rounded toward zero to whole shares (int() does this for both signs). It then sends that target minus the shares it holds, minus the shares in orders still waiting on the ticker, counting a waiting buy as positive and a waiting sell as negative.
0.2 * net_worth / price alone is the target holding. Sent as an order while the agent already holds the name, it buys the whole position again. A waiting limit order counts because it can still fill: a market order sent for the full gap while a bid waits can fill twice. A new tf.Limit on the same ticker replaces the waiting one, so when the agent resends a limit it sizes it against what is held alone.
The library accepts a fraction of a share, so rounding is the agent's choice. A whole number matches what a broker takes, and the hosted app refuses fractions. evaluate warns when every order in a step is below one share, which is the usual sign of weights sent as shares.
import tradefloor as tf
def shares_to_send(obs, ticker, weight):
"""Target holding, current holding, waiting orders and the order to send."""
# whole shares, rounded toward zero: floor() for a long target
target = int(weight * obs.portfolio.net_worth() / obs.price(ticker))
held = obs.position(ticker)
waiting = sum(o["remaining"] if o["side"] == "buy" else -o["remaining"]
for o in obs.portfolio.open_orders() if o["ticker"] == ticker)
return target, held, waiting, target - held - waiting
class TwentyPercent:
# hold 20% of net worth in AAA: bid for it at the first open,
# top up at the market when the gap passes 25 shares
def act(self, obs):
target, held, waiting, order = shares_to_send(obs, "AAA", 0.2)
if obs.step in (0, 1, 6, 7):
print(f"day {obs.day} step {obs.step_of_day}: target {target}, "
f"held {held:.0f}, waiting {waiting:.0f}, order {order:.0f}")
if obs.step == 0:
return {"AAA": tf.Limit(order, round(obs.price("AAA") * 0.997, 2))}
if obs.is_last_step_of_day:
return {"AAA": tf.Cancel()}
return {"AAA": order} if abs(order) > 25 else None
market = tf.Universe.random(40, seed=111)
card = tf.evaluate({"twenty": TwentyPercent()}, seed=7, universe=market, days=5)["twenty"]
print(card)day 0 step 0: target 997, held 0, waiting 0, order 997
day 0 step 1: target 987, held 0, waiting 997, order -10
day 1 step 0: target 962, held 0, waiting 0, order 962
day 1 step 1: target 968, held 962, waiting 0, order 6
Scorecard('twenty', pnl=-7,057, return=-0.71%, trades=2, impact=+0.97bps, sharpe=n/a (short run), vol=8.1%, in_market=80%, exposure=0.16x)On day 0 the bid 0.3% under the price waits all day and never fills. At step 1 the target is 987 shares and the bid still covers 997 of them, so the order is −10 and inside the band, and the agent sends nothing. Without the waiting term it would have bought 987 shares at the market on top of the bid. The day's last step cancels the bid, so on day 1 nothing waits and the agent buys its 962 shares at the market. One step later it holds them, and the gap is 6 shares. Shares and portfolio weights covers the usual sizing mistakes.
The agent's view
The observation holds what a trader inside the market could know: the roster and last prices, each name's order book and average daily volume, the agent's own positions, cash and waiting orders, daily bars and the published economic figures in obs.history, and a view of the market in obs.engine. What obs.engine is depends on how the agent is run:
| Access | How it is granted | obs.engine and obs.portfolio | Scorecard says |
|---|---|---|---|
| Ordinary | The default | Read-only views. The market view serves prices, the public columns, each book, the published macro figures and the yield curve, and which names have news today. It has no fair value, no factor attribution, no future macro path and no fork. | nothing |
| Privileged | privileged = True on the agent | The same read-only views, plus obs.hidden: every column, the factor attribution, the model's coefficients, the current macro state and economy, and the fundamentals. Still read-only, with no fork and no future path. | uses_hidden_state |
| Trusted | trusted_agents=True on the call | The live engine and portfolio, as every run handed them out before 0.8.5. An agent can fork the engine and run the copy ahead, read the macro path a scenario will set, and write to the market. | trusted |
Asking the ordinary view for anything it does not serve raises tf.SandboxError, which evaluate records in errors. On pt-v20 the business cycle and GDP growth arrive late, as the agencies publish them. The Oracle baseline is privileged. Under any access the harness compares the engine before and after each call to act and scores an agent that changed or copied it tampered. That check catches accidents and is no security boundary: agent code runs in the harness's own process, so run code you did not write in a separate process. What the agent sees in the reference lists each view's members.
Warm-up history
Every run starts at day 0 with no price history, so a rule that needs 20 days of bars would sit in cash for 20 days. history_days=N runs the market for N days before day 0 with nobody trading, and obs.history holds those days at the first decision.
import tradefloor as tf
market = tf.Universe.random(40, seed=111)
class Breakout:
# once a day, buy 300 shares of a name trading above its 20-day high
def act(self, obs):
if not obs.is_first_step_of_day:
return None
orders = {}
for ticker in obs.tickers:
bars = obs.history.bars(ticker, last=20)
if len(bars) == 20 and obs.position(ticker) == 0 \
and obs.price(ticker) > max(bar["high"] for bar in bars):
orders[ticker] = 300
return orders
card = tf.evaluate({"breakout": Breakout()}, seed=7, universe=market,
days=30, history_days=20)["breakout"]
print(card)Scorecard('breakout', pnl=23,386, return=+2.34%, trades=15, impact=+0.06bps, sharpe=+3.93, vol=5.0%, in_market=97%, exposure=0.43x, history_days=20)The warm-up changes the market: the scored days continue the warmed one, so on the same seed they are different days from a run without it, and the scorecard records history_days. It takes at most 2,520 days, ten years of 252. Missing warm-up history covers an agent that sits in cash because it has none.
Next steps
- Compare strategies puts the agent beside buy-and-hold and ranks it across many seeds.
- LLM adapters runs an agent built in an LLM framework through the same
actmethod. - Agents and evaluation has the signatures, the scorecard's fields and how
evaluatetreats errors.