Vet a trading bot
For developers with a trading bot. Your bot buys a stock when it rises above its 20-day high and sells when it falls below its 10-day low or 8% under the price it paid. It looks good on paper, and you are about to put real money behind it. The question is whether it does better than the simplest alternative: buying every stock and holding them. If it is no better, you would be paying its trading costs and taking its risk for nothing. This script runs the bot and buy-and-hold side by side on sixteen markets and counts which did better on each.
Don't take it live yet. The bot beat buying and holding on 7 of 16 markets, about what a coin toss would give, so this run finds no edge. Its own trading also cost 1,052 on one market, about a seventh of its profit there.
The bot against buy-and-hold, market by market
Profit of the bot less profit of buy-and-hold, each market starting from 1,000,000. Bars to the right are markets the bot won.
Bot less buy-and-hold
Sixteen markets instead of one
If you flip a coin once and it lands heads, you have learned nothing about the coin. A backtest on one price history is one flip: the bot either looks clever or it doesn't, and you cannot tell skill from luck. This script runs the bot and buy-and-hold side by side on sixteen different markets and counts which did better on each.
The question
Does my breakout bot do better than buying everything and holding it, before I put real money behind it?
The fair test
Sixteen different markets, with the bot and buy-and-hold on the same one each time. The test counts who won each market, so one lucky market cannot carry the result.
The result
The bot won 7 and lost the rest. The test puts the chance of a split at least that lopsided by luck alone at 0.80, so the bot is not better. Its trading also cost about a seventh of its profit on one market.
tradefloor against a backtest
A backtest replays one price history, so it gives the bot one market, and it guesses what trading cost. tradefloor builds as many markets as you ask for, plays the bot and buy-and-hold in each, and fills the bot's orders against an order book, so the cost is measured rather than assumed.
| Needed for this decision | Backtest on price history | tradefloor |
|---|---|---|
| More than one market to test on | One history | Sixteen markets here, and more by changing one line |
| A fair comparison with buy-and-hold | Same history, but one sample | Same market each time, counted with a paired sign test |
| What the bot's own orders cost | A fixed slippage guess | Measured against the same market without the bot's orders |
| A result someone else can check | Depends on their data | The same seed gives the same market on any machine |
The script
Install with pip install tradefloor. The script needs no API key and runs in under a minute. history_days=20 gives each market twenty days of price history before day 0, so the bot's first decision already has a 20-day high to compare with.
import tradefloor as tf
market = tf.Universe.random(10, seed=111)
SEEDS = range(16)
DAYS = 60
STOP = 0.08 # sell if a holding falls 8% below its entry price
class Breakout:
# buy above the 20-day high, sell below the 10-day low or at the stop
def __init__(self):
self.entry = {}
def act(self, obs):
orders = {}
for ticker in obs.tickers:
# checked at every step, since there are no stop orders to rest
price, held = obs.price(ticker), obs.position(ticker)
bars = obs.history.bars(ticker, last=20)
if held > 0 and (price < self.entry[ticker] * (1 - STOP)
or price < min(bar["low"] for bar in bars[-10:])):
orders[ticker] = -held
elif held == 0 and price > max(bar["high"] for bar in bars):
# one equal slot per name, the size buy-and-hold holds
slot = obs.portfolio.net_worth() / len(obs.tickers)
orders[ticker] = int(slot / price)
self.entry[ticker] = price
return orders
def entrants():
return {"breakout": Breakout(), "buy_and_hold": tf.baselines.BuyAndHold()}
# 20 days of warm-up, so the first decision already has a 20-day high
ranking = tf.rank(entrants, seeds=SEEDS, universe=market, days=DAYS,
history_days=20, workers=4)
bot, bench = ranking.records["breakout"], ranking.records["buy_and_hold"]
test = ranking.separation("breakout", "buy_and_hold")
# per seed, in currency: the bot's P&L less buy-and-hold's
excess = {f"seed {s}": round(x) for s, x in zip(ranking.seeds, bot.excess_pnls)}
# seed 0 again, priced against the same market with the bot's orders left out
cost = tf.tca.analyse(Breakout(), seed=0, universe=market, days=DAYS,
history_days=20)
paid = cost.shortfall()
seed0 = {"breakout before costs": round(bot.pnls[0] + paid),
"breakout after costs": round(bot.pnls[0]),
"buy_and_hold": round(bench.pnls[0])}
n = len(ranking.seeds)
print(f"{n} markets, {DAYS} days each, {len(market)} companies, 1,000,000 to start")
print(f"{'':14}{'median P&L':>12}{'mean vs B&H':>13}{'ahead':>8}")
print(f"{'breakout':14}{bot.median_pnl:>+12,.0f}{bot.mean_excess_pnl:>+13,.0f}"
f"{bot.seeds_ahead:>5}/{n}")
print(f"{'buy_and_hold':14}{bench.median_pnl:>+12,.0f}")
print(f"seed 0: traded {cost.as_dict()['notional']:,.0f} in {len(cost.fills)} fills, "
f"costs {paid:,.0f} ({cost.shortfall_bps():.1f} bps), P&L "
f"{seed0['breakout before costs']:+,} before costs, "
f"{seed0['breakout after costs']:+,} after")
if test["p_value"] < 0.05:
better = "beats" if test["wins_a"] > test["wins_b"] else "trails"
verdict = f"the bot {better} buy-and-hold on these markets"
else:
verdict = "no clear edge over buy-and-hold, so this is no case to go live"
print(f"Verdict: ahead on {test['wins_a']} of {test['paired_seeds']} markets, "
f"p = {test['p_value']:.2f}: {verdict}.")16 markets, 60 days each, 10 companies, 1,000,000 to start
median P&L mean vs B&H ahead
breakout -1,433 -300 7/16
buy_and_hold +8,096
seed 0: traded 1,487,662 in 15 fills, costs 1,052 (7.1 bps), P&L +7,679 before costs, +6,626 after
Verdict: ahead on 7 of 16 markets, p = 0.80: no clear edge over buy-and-hold, so this is no case to go live.What costs took on one market
Profit on market 0, before and after the bot's trading costs
Currency, 60 days, with buy-and-hold on the same market
Market 0 fell, so the bot's lead there comes from sitting in cash, not from its timing. The bot is in cash much of the time and buy-and-hold never is, so the bot's gap on each market mostly follows whether that market rose or fell.
Adapting it
Replace the body of Breakout.act with your own rule. Raise SEEDS to detect a smaller edge, or DAYS for longer markets. The strategy comparison guide explains the sign test and what tf.rank reports.
Limits
- Ten made-up companies on the default preset, and 60 trading days per market. Longer runs or another roster may differ.
- Profit only, with no adjustment for risk. The bot carries far less exposure than buy-and-hold.
- The cost line is for one market, not an average over all sixteen.
- Equities and cash only, and no stop orders resting in the book. The stop is checked at each decision step, so a price can fall past it between checks.
- Sixteen markets cannot detect a small edge. A real but small difference needs many more.
Reference
tf.rank, Ranking.separation and tf.tca.analyse are in Agents and evaluation.