A/B test an AI agent
For teams building AI trading agents. Version 1 is the agent you run today. Version 2 adds one line to its prompt: hold less of volatile stocks, and halve every position when markets panic. The worry is that the extra caution costs money in normal times. Ship a change that quietly loses money and every user pays for it. Skip one that protects them, and the next crash does. This script tests both versions on the same markets, ordinary ones and ones in a crisis, so you can decide whether to ship the change.
Ship version 2. In ordinary markets the extra line made no difference. When a crisis hit, version 2 lost 120,752 more than in the same market without it. Version 1 lost 252,203 more, about twice as much.
The same market, calm and in a crisis
Net worth from 1,000,000. On day 10 the market is copied and one copy goes into a crisis. Each version trades both copies.
Version 2 in the crisis880,513Version 1 in the crisis750,156Version 2 calm1,001,265Version 1 calm1,002,359
Both versions on the same markets
Car makers crash-test two designs into the same wall at the same speed. If each car hit a different wall, a better result could just mean a softer wall. Testing an agent on one stretch of history is like that, and the history may not even have a crisis in it. This script tests both versions on the same forty markets, half of them with a crisis.
The question
Does the risk rule in version 2 cost money in normal times, and does it save money in a crisis?
The fair test
Both versions trade the same twenty ordinary markets and the same twenty crisis markets. Then one market is copied on day 10 and the crisis hits both versions in the same copy.
The result
In ordinary markets the rule made no difference that could be measured. In a crisis version 2 lost about half as much, and led on 15 of 20 markets. Ship version 2.
tradefloor against a backtest
A backtest replays one history, so both versions get one sample, and the crisis you care about may not be in it. tradefloor can put the crisis where you need it, run both versions through it on many markets, and copy one market so both meet the crisis at the same moment.
| Needed for this decision | Backtest on price history | tradefloor |
|---|---|---|
| A crisis to test against | Only if the dates happen to include one | Any shipped scenario, started on the day you choose |
| Both versions under the same conditions | Same history, but one sample | Forty paired markets and a sign test |
| The crisis landing at the same moment for both | No | Yes. A fork copies the market, and the crisis lands on both copies. |
| Model calls that cost nothing to repeat | Every rerun calls the model again | Record the answers once and replay them for free |
The script
Install with pip install tradefloor. The script needs no API key: a plain function stands in for each model call, with the comment saying what each version's prompt asks for. It runs in under a minute.
import json
import tradefloor as tf
from tradefloor.integrations.callable import callable_agent
market = tf.Universe.random(12, seed=11)
crisis = tf.Scenario.load("liquidity_crisis")
def decide(payload, careful):
# stands in for the model's answer: v1's prompt says "buy the five strongest movers",
# v2's adds "smaller in volatile names, and halve everything when the VIX is over 30"
assets = [a for a in payload["assets"] if a["return_5d"] is not None]
picks = sorted(assets, key=lambda a: a["return_5d"], reverse=True)[:5]
typical_vol = sorted(a["volatility"] for a in assets)[len(assets) // 2] if assets else 0
actions = []
for a in payload["assets"]:
weight = 0.18 if a in picks else 0.0
if careful and weight:
weight *= min(1.0, typical_vol / a["volatility"])
weight /= 2 if payload["macro"]["vix"] > 30 else 1
gap = int(weight * payload["portfolio"]["net_worth"] / a["price"] - a["position"])
if gap:
actions.append({"symbol": a["symbol"], "side": "BUY" if gap > 0 else "SELL",
"quantity": abs(gap)})
return json.dumps({"actions": actions, "rationale": "careful" if careful else "chase"})
def v1(payload):
return decide(payload, careful=False)
def v2(payload):
return decide(payload, careful=True)
def both():
# new adapters for every market, so nothing carries from one seed to the next
return {"v1": callable_agent(v1), "v2": callable_agent(v2)}
# Paired A/B tests: both versions meet the same markets, ordinary and in a crisis
tests = {
"ordinary, 20 days": tf.rank(both, seeds=range(20), universe=market, days=20,
benchmark="v1", workers=8),
"crisis, 16 days": tf.rank(both, seeds=range(20), universe=market, days=16,
benchmark="v1", scenario=crisis.starting_at(6)),
}
pnl_gap = {kind: [round(g) for g in r.records["v2"].excess_pnls] for kind, r in tests.items()}
# One shared past, then the same crisis for both versions from day 10
net_worth = {}
for name, fn in (("v1", v1), ("v2", v2)):
world = tf.World(seed=7, universe=market, agent=callable_agent(fn))
path = []
for day in range(10):
world.run(1)
path.append(round(world.net_worth()))
calm, stressed = world.fork("calm", "crisis")
stressed.apply(crisis, at=0)
assert tf.agree(calm, stressed).identical
net_worth[f"{name} calm"], net_worth[f"{name} crisis"] = list(path), list(path)
for day in range(20):
for arm, label in ((calm, "calm"), (stressed, "crisis")):
arm.run(1)
net_worth[f"{name} {label}"].append(round(arm.net_worth()))
print("Paired test, v2 against v1 on the same markets")
print(f"{'markets':18}{'count':>6}{'v2 ahead':>10}{'p':>7}{'mean v2 - v1':>14}")
for kind, r in tests.items():
s = r.separation("v2", "v1")
print(f"{kind:18}{s['paired_seeds']:>6}{s['wins_a']:>10}{s['p_value']:>7.2f}"
f"{r.records['v2'].mean_excess_pnl:>+14,.0f}")
print()
print("One market forked on day 10, then 20 days calm or in crisis")
print(f"{'version':8}{'calm end':>12}{'crisis end':>12}{'crisis cost':>13}")
crisis_loss = {}
for name in ("v1", "v2"):
calm_end, crisis_end = net_worth[f"{name} calm"][-1], net_worth[f"{name} crisis"][-1]
crisis_loss[name] = calm_end - crisis_end
print(f"{name:8}{calm_end:>12,}{crisis_end:>12,}{-crisis_loss[name]:>+13,}")
ordinary, stress = (r.separation("v2", "v1") for r in tests.values())
worse_day_to_day = ordinary["p_value"] < 0.05 and ordinary["wins_b"] > ordinary["wins_a"]
better_in_crisis = stress["p_value"] < 0.05 and stress["wins_a"] > stress["wins_b"]
decision = "ship v2" if better_in_crisis and not worse_day_to_day else "not yet, add seeds"
print(f"Verdict: {decision}. v2 was ahead on {ordinary['wins_a']} of {ordinary['paired_seeds']} "
f"ordinary markets (p = {ordinary['p_value']:.2f}) and {stress['wins_a']} of "
f"{stress['paired_seeds']} crisis markets (p = {stress['p_value']:.2f}).")Paired test, v2 against v1 on the same markets markets count v2 ahead p mean v2 - v1 ordinary, 20 days 20 11 0.82 -207 crisis, 16 days 20 15 0.04 +49,477 One market forked on day 10, then 20 days calm or in crisis version calm end crisis end crisis cost v1 1,002,359 750,156 -252,203 v2 1,001,265 880,513 -120,752 Verdict: ship v2. v2 was ahead on 11 of 20 ordinary markets (p = 0.82) and 15 of 20 crisis markets (p = 0.04).
The p column is the chance of a split at least that lopsided if the two versions were equally good. At 0.82 the ordinary markets show no difference. At 0.04 the crisis markets show a real one, though close to the usual 0.05 line.
Adapting it
Replace the bodies of v1 and v2 with calls to your model under each prompt, or pass two framework adapters into both(); the LLM adapters page lists them. Change liquidity_crisis to the scenario you care about. Record a live run with Record and replay so the comparison can be rerun without paying for the model again.
Limits
- The crisis result is close to the line: p = 0.04 on twenty markets. Run more crisis markets before quoting the size of the gap.
- No measurable difference in ordinary markets is not proof of none. Twenty short markets cannot see a small cost.
- One scenario, started early so the short runs reach it. Its shocks are stated assumptions, not a forecast, and other scenarios may rank the versions differently.
- Both versions are fixed functions standing in for model calls. A real AI agent answers differently from run to run, which is what recording is for.
- Twelve simulated companies on the default preset.
Reference
tf.rank and Ranking.separation are in Agents and evaluation. The callable adapter is in LLM adapters, and forks are in Fork a market.