Skip to the page
GUIDES/WRITE AN AGENT

Write an agent

An agent is any Python object with an act(obs) method. This guide builds one, sizes its orders, reads what it is allowed to see and gives it price history before its first decision. Every argument, default and return value is in the reference, Agents and evaluation.

Every code block on this page runs on its own.

A first agent

At each decision step, six a day by default, the harness builds an observation, passes it to act and reads back a dict of ticker to order. tf.evaluate runs the agent on a market built from a seed and a universe, here 40 generated companies, and returns a scorecard per agent.

import tradefloor as tf

market = tf.Universe.random(40, seed=111)

class Quoter:
    # bid a cent under the last price at each open, cancel what is left at the close
    def act(self, obs):
        if obs.is_first_step_of_day:
            bid = round(obs.price("AAA") - 0.01, 2)
            return {"AAA": tf.Limit(500, bid)}
        if obs.is_last_step_of_day:
            return {"AAA": tf.Cancel()}
        return None

card = tf.evaluate({"quoter": Quoter()}, seed=7, universe=market, days=10)["quoter"]
print(card)
Scorecard('quoter', pnl=20,597, return=+2.06%, trades=39, impact=+1.47bps, sharpe=n/a (short run), vol=18.4%, in_market=100%, exposure=0.56x)

trades counts fills. Ten limit orders made 39 of them, because a waiting order fills in parts as the price comes to it. Each agent starts with 1,000,000 in cash and may hold positions worth up to twice its net worth.

obs.step counts decision steps across the whole run, so day 1 starts at step 6. A guard written if obs.step == 0 therefore fires once a run, and an agent that means once a day reads obs.is_first_step_of_day or obs.step_of_day instead.

Orders from act

act returns a dict of ticker to order, or None to trade nothing that step. A plain number is a market order for that many shares, positive to buy and negative to sell or short. tf.Limit(quantity, price) waits in the book at its price or better, and tf.Cancel() withdraws every waiting order on the ticker. A new tf.Limit on a ticker replaces the one waiting there. The full table, with what happens to a partial fill, is in What act returns.

A bad entry is refused on its own and the rest of the dict still trades. Each refusal adds a line to the scorecard's errors, so read errors before the P&L:

import tradefloor as tf

market = tf.Universe.random(40, seed=111)

class Careless:
    # a quantity sent as text, and a ticker the market does not list
    def act(self, obs):
        if obs.step == 0:
            return {"AAA": 100, "AAB": "100", "ZZZZ": 50}
        return None

card = tf.evaluate({"careless": Careless()}, seed=7, universe=market, days=1)["careless"]
print(card.trades, card.rejected)
for line in card.errors:
    print(line)
1 2
step 0: the order for 'AAB' must be a number of shares, a tf.Limit or a tf.Cancel, got '100' (str)
step 0: no instrument with ticker "ZZZZ" in this universe

If an agent's orders never fill or are refused, An agent that makes no trades and Rejected orders list the usual causes.

Sizing to a target weight

Values in the dict are shares, never weights. To hold 20% of net worth in AAA, the agent works out the holding it wants, 0.2 × net worth ÷ price rounded toward zero to whole shares (int() does this for both signs). It then sends that target minus the shares it holds, minus the shares in orders still waiting on the ticker, counting a waiting buy as positive and a waiting sell as negative.

0.2 * net_worth / price alone is the target holding. Sent as an order while the agent already holds the name, it buys the whole position again. A waiting limit order counts because it can still fill: a market order sent for the full gap while a bid waits can fill twice. A new tf.Limit on the same ticker replaces the waiting one, so when the agent resends a limit it sizes it against what is held alone.

The library accepts a fraction of a share, so rounding is the agent's choice. A whole number matches what a broker takes, and the hosted app refuses fractions. evaluate warns when every order in a step is below one share, which is the usual sign of weights sent as shares.

import tradefloor as tf

def shares_to_send(obs, ticker, weight):
    """Target holding, current holding, waiting orders and the order to send."""
    # whole shares, rounded toward zero: floor() for a long target
    target = int(weight * obs.portfolio.net_worth() / obs.price(ticker))
    held = obs.position(ticker)
    waiting = sum(o["remaining"] if o["side"] == "buy" else -o["remaining"]
                  for o in obs.portfolio.open_orders() if o["ticker"] == ticker)
    return target, held, waiting, target - held - waiting

class TwentyPercent:
    # hold 20% of net worth in AAA: bid for it at the first open,
    # top up at the market when the gap passes 25 shares
    def act(self, obs):
        target, held, waiting, order = shares_to_send(obs, "AAA", 0.2)
        if obs.step in (0, 1, 6, 7):
            print(f"day {obs.day} step {obs.step_of_day}: target {target}, "
                  f"held {held:.0f}, waiting {waiting:.0f}, order {order:.0f}")
        if obs.step == 0:
            return {"AAA": tf.Limit(order, round(obs.price("AAA") * 0.997, 2))}
        if obs.is_last_step_of_day:
            return {"AAA": tf.Cancel()}
        return {"AAA": order} if abs(order) > 25 else None

market = tf.Universe.random(40, seed=111)
card = tf.evaluate({"twenty": TwentyPercent()}, seed=7, universe=market, days=5)["twenty"]
print(card)
day 0 step 0: target 997, held 0, waiting 0, order 997
day 0 step 1: target 987, held 0, waiting 997, order -10
day 1 step 0: target 962, held 0, waiting 0, order 962
day 1 step 1: target 968, held 962, waiting 0, order 6
Scorecard('twenty', pnl=-7,057, return=-0.71%, trades=2, impact=+0.97bps, sharpe=n/a (short run), vol=8.1%, in_market=80%, exposure=0.16x)

On day 0 the bid 0.3% under the price waits all day and never fills. At step 1 the target is 987 shares and the bid still covers 997 of them, so the order is −10 and inside the band, and the agent sends nothing. Without the waiting term it would have bought 987 shares at the market on top of the bid. The day's last step cancels the bid, so on day 1 nothing waits and the agent buys its 962 shares at the market. One step later it holds them, and the gap is 6 shares. Shares and portfolio weights covers the usual sizing mistakes.

The agent's view

The observation holds what a trader inside the market could know: the roster and last prices, each name's order book and average daily volume, the agent's own positions, cash and waiting orders, daily bars and the published economic figures in obs.history, and a view of the market in obs.engine. What obs.engine is depends on how the agent is run:

AccessHow it is grantedobs.engine and obs.portfolioScorecard says
OrdinaryThe defaultRead-only views. The market view serves prices, the public columns, each book, the published macro figures and the yield curve, and which names have news today. It has no fair value, no factor attribution, no future macro path and no fork.nothing
Privilegedprivileged = True on the agentThe same read-only views, plus obs.hidden: every column, the factor attribution, the model's coefficients, the current macro state and economy, and the fundamentals. Still read-only, with no fork and no future path.uses_hidden_state
Trustedtrusted_agents=True on the callThe live engine and portfolio, as every run handed them out before 0.8.5. An agent can fork the engine and run the copy ahead, read the macro path a scenario will set, and write to the market.trusted

Asking the ordinary view for anything it does not serve raises tf.SandboxError, which evaluate records in errors. On pt-v20 the business cycle and GDP growth arrive late, as the agencies publish them. The Oracle baseline is privileged. Under any access the harness compares the engine before and after each call to act and scores an agent that changed or copied it tampered. That check catches accidents and is no security boundary: agent code runs in the harness's own process, so run code you did not write in a separate process. What the agent sees in the reference lists each view's members.

Warm-up history

Every run starts at day 0 with no price history, so a rule that needs 20 days of bars would sit in cash for 20 days. history_days=N runs the market for N days before day 0 with nobody trading, and obs.history holds those days at the first decision.

import tradefloor as tf

market = tf.Universe.random(40, seed=111)

class Breakout:
    # once a day, buy 300 shares of a name trading above its 20-day high
    def act(self, obs):
        if not obs.is_first_step_of_day:
            return None
        orders = {}
        for ticker in obs.tickers:
            bars = obs.history.bars(ticker, last=20)
            if len(bars) == 20 and obs.position(ticker) == 0 \
                    and obs.price(ticker) > max(bar["high"] for bar in bars):
                orders[ticker] = 300
        return orders

card = tf.evaluate({"breakout": Breakout()}, seed=7, universe=market,
                   days=30, history_days=20)["breakout"]
print(card)
Scorecard('breakout', pnl=23,386, return=+2.34%, trades=15, impact=+0.06bps, sharpe=+3.93, vol=5.0%, in_market=97%, exposure=0.43x, history_days=20)

The warm-up changes the market: the scored days continue the warmed one, so on the same seed they are different days from a run without it, and the scorecard records history_days. It takes at most 2,520 days, ten years of 252. Missing warm-up history covers an agent that sits in cash because it has none.

Next steps

  • Compare strategies puts the agent beside buy-and-hold and ranks it across many seeds.
  • LLM adapters runs an agent built in an LLM framework through the same act method.
  • Agents and evaluation has the signatures, the scorecard's fields and how evaluate treats errors.