Skip to the page
EXAMPLES/A/B TEST AN AI AGENT

A/B test an AI agent

For teams building AI trading agents. Version 1 is the agent you run today. Version 2 adds one line to its prompt: hold less of volatile stocks, and halve every position when markets panic. The worry is that the extra caution costs money in normal times. Ship a change that quietly loses money and every user pays for it. Skip one that protects them, and the next crash does. This script tests both versions on the same markets, ordinary ones and ones in a crisis, so you can decide whether to ship the change.

Ship version 2. In ordinary markets the extra line made no difference. When a crisis hit, version 2 lost 120,752 more than in the same market without it. Version 1 lost 252,203 more, about twice as much.

120,752crisis loss, version 2
252,203crisis loss, version 1
15 of 20crisis markets where version 2 did better
11 of 20ordinary markets where it did better, about even

Both versions on the same markets

Car makers crash-test two designs into the same wall at the same speed. If each car hit a different wall, a better result could just mean a softer wall. Testing an agent on one stretch of history is like that, and the history may not even have a crisis in it. This script tests both versions on the same forty markets, half of them with a crisis.

The question

Does the risk rule in version 2 cost money in normal times, and does it save money in a crisis?

The fair test

Both versions trade the same twenty ordinary markets and the same twenty crisis markets. Then one market is copied on day 10 and the crisis hits both versions in the same copy.

The result

In ordinary markets the rule made no difference that could be measured. In a crisis version 2 lost about half as much, and led on 15 of 20 markets. Ship version 2.

tradefloor against a backtest

A backtest replays one history, so both versions get one sample, and the crisis you care about may not be in it. tradefloor can put the crisis where you need it, run both versions through it on many markets, and copy one market so both meet the crisis at the same moment.

Needed for this decisionBacktest on price historytradefloor
A crisis to test againstOnly if the dates happen to include oneAny shipped scenario, started on the day you choose
Both versions under the same conditionsSame history, but one sampleForty paired markets and a sign test
The crisis landing at the same moment for bothNoYes. A fork copies the market, and the crisis lands on both copies.
Model calls that cost nothing to repeatEvery rerun calls the model againRecord the answers once and replay them for free

The script

Install with pip install tradefloor. The script needs no API key: a plain function stands in for each model call, with the comment saying what each version's prompt asks for. It runs in under a minute.

import json

import tradefloor as tf
from tradefloor.integrations.callable import callable_agent

market = tf.Universe.random(12, seed=11)
crisis = tf.Scenario.load("liquidity_crisis")


def decide(payload, careful):
    # stands in for the model's answer: v1's prompt says "buy the five strongest movers",
    # v2's adds "smaller in volatile names, and halve everything when the VIX is over 30"
    assets = [a for a in payload["assets"] if a["return_5d"] is not None]
    picks = sorted(assets, key=lambda a: a["return_5d"], reverse=True)[:5]
    typical_vol = sorted(a["volatility"] for a in assets)[len(assets) // 2] if assets else 0
    actions = []
    for a in payload["assets"]:
        weight = 0.18 if a in picks else 0.0
        if careful and weight:
            weight *= min(1.0, typical_vol / a["volatility"])
            weight /= 2 if payload["macro"]["vix"] > 30 else 1
        gap = int(weight * payload["portfolio"]["net_worth"] / a["price"] - a["position"])
        if gap:
            actions.append({"symbol": a["symbol"], "side": "BUY" if gap > 0 else "SELL",
                            "quantity": abs(gap)})
    return json.dumps({"actions": actions, "rationale": "careful" if careful else "chase"})


def v1(payload):
    return decide(payload, careful=False)


def v2(payload):
    return decide(payload, careful=True)


def both():
    # new adapters for every market, so nothing carries from one seed to the next
    return {"v1": callable_agent(v1), "v2": callable_agent(v2)}


# Paired A/B tests: both versions meet the same markets, ordinary and in a crisis
tests = {
    "ordinary, 20 days": tf.rank(both, seeds=range(20), universe=market, days=20,
                                 benchmark="v1", workers=8),
    "crisis, 16 days": tf.rank(both, seeds=range(20), universe=market, days=16,
                               benchmark="v1", scenario=crisis.starting_at(6)),
}
pnl_gap = {kind: [round(g) for g in r.records["v2"].excess_pnls] for kind, r in tests.items()}

# One shared past, then the same crisis for both versions from day 10
net_worth = {}
for name, fn in (("v1", v1), ("v2", v2)):
    world = tf.World(seed=7, universe=market, agent=callable_agent(fn))
    path = []
    for day in range(10):
        world.run(1)
        path.append(round(world.net_worth()))
    calm, stressed = world.fork("calm", "crisis")
    stressed.apply(crisis, at=0)
    assert tf.agree(calm, stressed).identical
    net_worth[f"{name} calm"], net_worth[f"{name} crisis"] = list(path), list(path)
    for day in range(20):
        for arm, label in ((calm, "calm"), (stressed, "crisis")):
            arm.run(1)
            net_worth[f"{name} {label}"].append(round(arm.net_worth()))

print("Paired test, v2 against v1 on the same markets")
print(f"{'markets':18}{'count':>6}{'v2 ahead':>10}{'p':>7}{'mean v2 - v1':>14}")
for kind, r in tests.items():
    s = r.separation("v2", "v1")
    print(f"{kind:18}{s['paired_seeds']:>6}{s['wins_a']:>10}{s['p_value']:>7.2f}"
          f"{r.records['v2'].mean_excess_pnl:>+14,.0f}")
print()
print("One market forked on day 10, then 20 days calm or in crisis")
print(f"{'version':8}{'calm end':>12}{'crisis end':>12}{'crisis cost':>13}")
crisis_loss = {}
for name in ("v1", "v2"):
    calm_end, crisis_end = net_worth[f"{name} calm"][-1], net_worth[f"{name} crisis"][-1]
    crisis_loss[name] = calm_end - crisis_end
    print(f"{name:8}{calm_end:>12,}{crisis_end:>12,}{-crisis_loss[name]:>+13,}")

ordinary, stress = (r.separation("v2", "v1") for r in tests.values())
worse_day_to_day = ordinary["p_value"] < 0.05 and ordinary["wins_b"] > ordinary["wins_a"]
better_in_crisis = stress["p_value"] < 0.05 and stress["wins_a"] > stress["wins_b"]
decision = "ship v2" if better_in_crisis and not worse_day_to_day else "not yet, add seeds"
print(f"Verdict: {decision}. v2 was ahead on {ordinary['wins_a']} of {ordinary['paired_seeds']} "
      f"ordinary markets (p = {ordinary['p_value']:.2f}) and {stress['wins_a']} of "
      f"{stress['paired_seeds']} crisis markets (p = {stress['p_value']:.2f}).")
Paired test, v2 against v1 on the same markets
markets            count  v2 ahead      p  mean v2 - v1
ordinary, 20 days     20        11   0.82          -207
crisis, 16 days       20        15   0.04       +49,477

One market forked on day 10, then 20 days calm or in crisis
version     calm end  crisis end  crisis cost
v1         1,002,359     750,156     -252,203
v2         1,001,265     880,513     -120,752
Verdict: ship v2. v2 was ahead on 11 of 20 ordinary markets (p = 0.82) and 15 of 20 crisis markets (p = 0.04).

The p column is the chance of a split at least that lopsided if the two versions were equally good. At 0.82 the ordinary markets show no difference. At 0.04 the crisis markets show a real one, though close to the usual 0.05 line.

Adapting it

Replace the bodies of v1 and v2 with calls to your model under each prompt, or pass two framework adapters into both(); the LLM adapters page lists them. Change liquidity_crisis to the scenario you care about. Record a live run with Record and replay so the comparison can be rerun without paying for the model again.

Limits

  • The crisis result is close to the line: p = 0.04 on twenty markets. Run more crisis markets before quoting the size of the gap.
  • No measurable difference in ordinary markets is not proof of none. Twenty short markets cannot see a small cost.
  • One scenario, started early so the short runs reach it. Its shocks are stated assumptions, not a forecast, and other scenarios may rank the versions differently.
  • Both versions are fixed functions standing in for model calls. A real AI agent answers differently from run to run, which is what recording is for.
  • Twelve simulated companies on the default preset.

Reference

tf.rank and Ranking.separation are in Agents and evaluation. The callable adapter is in LLM adapters, and forks are in Fork a market.