Skip to the page
GUIDES/COMPARE STRATEGIES

Compare strategies

One market is one sample. This guide puts an agent beside buy-and-hold on a single market, then on many markets with tf.rank, reads the paired result, and prices what its own trading cost. The signatures and every field of the results are in the reference, Agents and evaluation.

Every code block on this page runs on its own. The agent they compare is the Quoter from Write an agent, repeated in each block, on a universe of 40 generated companies.

One market and the baselines

tf.reference_agents() builds five baselines: buy-and-hold, a random trader, momentum, mean-reversion and the Oracle, which reads the model's fair value. tf.evaluate gives each agent its own copy of the same market, so no agent uses up liquidity another expected, and tf.versus_buy_and_hold gives each agent's P&L minus buy-and-hold's.

import tradefloor as tf

market = tf.Universe.random(40, seed=111)

class Quoter:
    # bid a cent under the last price at each open, cancel what is left at the close
    def act(self, obs):
        if obs.is_first_step_of_day:
            bid = round(obs.price("AAA") - 0.01, 2)
            return {"AAA": tf.Limit(500, bid)}
        if obs.is_last_step_of_day:
            return {"AAA": tf.Cancel()}
        return None

agents = tf.reference_agents(seed=3)
agents["quoter"] = Quoter()
scores = tf.evaluate(agents, seed=7, universe=market, days=10)
for name, gap in tf.versus_buy_and_hold(scores).items():
    print(f"{name:15} {gap:+12,.0f}")
random               -41,359
momentum             -36,407
mean_reversion       -44,998
oracle                -7,242
quoter               +10,121

On pt-v20 compare with buy-and-hold. A shock there moves fair value for good, so the Oracle's P&L follows the market's month and sets no ceiling, and tf.capture_ratio returns nothing and says why. To put several agents in one market, where one agent's trades move the prices another meets, use a World.

Across seeds with rank

tf.rank runs the same agents on many seeds, each a different market on the same universe. It calls the factory once per seed, so every seed gets new agents and no state carries from one market to the next.

import tradefloor as tf

market = tf.Universe.random(40, seed=111)

class Quoter:
    # bid a cent under the last price at each open, cancel what is left at the close
    def act(self, obs):
        if obs.is_first_step_of_day:
            bid = round(obs.price("AAA") - 0.01, 2)
            return {"AAA": tf.Limit(500, bid)}
        if obs.is_last_step_of_day:
            return {"AAA": tf.Cancel()}
        return None

def entrants():
    # new agents for every seed, so no state carries from one market to the next
    agents = tf.reference_agents(seed=3)
    agents["quoter"] = Quoter()
    return agents

ranking = tf.rank(entrants, seeds=range(8), universe=market, days=10)
print(ranking.report())
print(ranking.separation("quoter", "buy_and_hold"))
8 seeds on universe 9be68b9bc37e... under model pt-v20
  quoter            vs buy_and_hold       +7,064 a seed  ahead 5/8  wins 5/8
  buy_and_hold      the benchmark  median pnl       +1,709  wins 2/8
  mean_reversion    vs buy_and_hold      -23,726 a seed  ahead 1/8  wins 1/8
  random            vs buy_and_hold      -30,372 a seed  ahead 0/8  wins 0/8
  momentum          vs buy_and_hold      -48,869 a seed  ahead 0/8  wins 0/8
  The Oracle 'oracle' ran and has no row, because it is not a ceiling on this model (below). Its P&L over buy_and_hold's was -595 a seed, ahead on 5 of 8 seeds.
  No capture ratio on pt-v20. Market moves there mostly stick: each shock moves fair value for good, so even perfect knowledge of the model's fair value leaves little edge. The Oracle made money in 10 of 14 test markets and its P&L follows the market's month, so a fraction of it would measure the month, not the agent. Compare against buy-and-hold instead.
{'a': 'quoter', 'b': 'buy_and_hold', 'wins_a': 5, 'wins_b': 3, 'ties': 0, 'paired_seeds': 8, 'decisive': False, 'p_value': 0.7265625}

Reading the result

Each row of the report is one agent against the benchmark, buy-and-hold by default, measured seed by seed on the same market. The figure before "a seed" is the mean of the agent's P&L minus buy-and-hold's, in currency, and on pt-v20 it is the figure to quote. "ahead" counts the seeds on which the agent earned more than buy-and-hold, and "wins" the seeds on which it had the highest P&L of every ranked agent.

separation(a, b) is a paired sign test between two agents. The Quoter was ahead on 5 of 8 seeds and behind on 3. At 5 wins in 8 and p = 0.73, this run can't tell them apart, because a fair coin lands 5 to 3 or further from even that often. decisive is true only when one agent won on every seed, and it is a count, never a significance test. Report the paired count beside any pooled P&L, and add seeds before calling a winner. ranking.records["quoter"] holds the per-seed figures behind the row, described in AgentRecord.

A row can also carry REFUSED, UNUSABLE or RAISED lines, which count the steps an agent's orders were refused, its act returned something other than a dict, or its code raised. Read those before the P&L. Ranking lists every member of the result.

Execution cost

tf.tca.analyse runs the same seed twice, with and without the agent's orders, and prices every fill against the market where it never traded. Real data cannot supply that untraded market, which is why arrival price and VWAP stand in for it there.

import tradefloor as tf

market = tf.Universe.random(40, seed=111)

class RoundTrip:
    # buy 2,000 AAA, sell them back three steps later
    def act(self, obs):
        if obs.step == 0:
            return {"AAA": 2000}
        if obs.step == 3:
            return {"AAA": -2000}
        return None

ex = tf.tca.analyse(RoundTrip(), seed=42, universe=market, days=1)
print(round(ex.shortfall_bps(), 2))
8.41

A positive shortfall is a cost. The engine has no slippage formula: a large order walks the book and pays for the depth it takes, and the cost of size follows the square-root law real metaorders follow (How it is measured, row C9). tca.analyse takes market orders only and refuses a tf.Limit, because the part of a limit order that waits fills inside a session, where the untraded run has no price to compare it with. tca.analyse and Execution have the arguments and the result.

Next steps

  • Fork a market changes one input in a copy of a market and compares the two futures.
  • Record and replay makes a comparison that involves a model reproducible.