Skip to the page
API REFERENCE/AGENTS AND EVALUATION

Agents and evaluation

Reference for the agent protocol, tf.evaluate, tf.rank, the baselines, tf.tca.analyse and the records they return: Scorecard, Ranking, AgentRecord, Execution and RunManifest. Worked examples are in the guides: Write an agent builds an agent and sizes its orders, Compare strategies ranks agents across seeds and prices their trading, and Record and replay publishes a run a reader can check.

Writing an agent

An agent is any Python object with an act(obs) method. At each decision step the harness passes it an observation and reads back orders. By default a run has 6 decision steps a day, 65 simulated minutes apart, and each agent starts with 1,000,000 in cash and may hold positions worth up to twice its net worth. obs.step counts steps across the whole run, and obs.step_of_day, obs.is_first_step_of_day and obs.is_last_step_of_day count within the day. An agent may also declare privileged = True and keep an explain(day) method, both described below. A first agent builds and scores one on a seeded market.

What act returns

A dict of ticker to order, where each value is one of these:

ValueOrder
1000 or -500A market order for that many shares. It fills against the book at once, and what the book cannot fill is listed in partial_fills.
tf.Limit(500, 101.25)Buy up to 500 shares at 101.25 or better, or sell with a negative quantity. What does not fill waits in the book, behind the shares already at that price, until it fills, you cancel it, or a new limit on the ticker replaces it.
tf.Cancel()Cancel every order of yours still waiting on that ticker.

Return None or {} to trade nothing. Values are shares, never weights, and any finite number of shares is accepted, fractions included. To hold a target weight, send the target holding less the shares already held and the shares in orders still waiting on the ticker, with the target rounded to whole shares if you want them:

target = int(weight * obs.portfolio.net_worth() / obs.price(ticker))
order  = target - obs.position(ticker) - waiting

The target alone, such as 0.2 * net_worth / price, is the holding to reach, and sent as an order while a position exists it buys the position again. Sizing to a target weight works through an example with a waiting order.

There are no stop or bracket orders, so check a stop yourself at each step. A refused entry, such as a string or NaN, gets a line in the scorecard's errors and the rest of the dict still trades, so read errors before the P&L.

What the agent sees

The observation holds the roster and last prices, each name's order book and average volume, the agent's own positions, cash and waiting orders, daily bars and the published economic figures in obs.history, and a view of the market in obs.engine. What obs.engine and obs.portfolio serve depends on the agent's access: ordinary by default, privileged with privileged = True on the agent, or trusted with trusted_agents=True on the call. The levels are the same under evaluate, rank, World and tca.analyse, and The agent's view sets out what each one serves.

Ordinary and privileged agents get the read-only MarketView and PortfolioView, which raise tf.SandboxError for anything they do not serve, recorded in errors. obs.hidden is None unless the agent is privileged, when it is a read-only HiddenState that also carries the size of each day's news. obs.engine.bars() refuses under evaluate, which records no days, so daily bars come from obs.history. An LLM framework run through an adapter sees only the payload the adapter builds, the same at every level: What the model sees.

Warm-up history

history_days=N on evaluate, rank, World or tca.analyse runs the market for N days before day 0 with nobody trading, and obs.history holds those days at the first decision, labeled -N to -1. Each scored day is added after its last step. N is at most 2,520, ten years of 252 days, and no scenario applies during the warm-up. Warm-up history runs a breakout rule on it.

Against buy-and-hold, across seeds

evaluate scores one market, which is one sample. Compare each agent with buy-and-hold on the same market with versus_buy_and_hold, and run rank over many seeds before calling a winner. Compare strategies does both and reads the paired sign test. To put several agents in one market, use a World.

Explaining moves

An agent may also have explain(day). The harness calls it after each close and expects the name of the factor that moved prices most that day. On pt-v20 random_noise wins almost every day, so the scorecard reports explanation_baseline, what naming the most common winner every day scored, and explanation_edge, the accuracy less that baseline, which is the figure to quote.

Execution cost

tf.tca.analyse prices every fill against the same seed run without the agent's orders. Execution cost in Compare strategies prices a round trip and says what the figure measures, and tca.analyse below has the arguments.

evaluate

Score each agent on its own copy of the same market.

tf.evaluate(
    agents: dict[str, Agent | StrategySpec],
    *, seed: int,
    universe: Sequence[Instrument],
    macro: Macro | None = None,
    days: int = 5,
    steps_per_day: int = 6,
    ticks_per_step: int = 65,
    cash: float = 1000000.0,
    max_leverage: float | None = 2.0,
    start: tuple[int, int, int] = (9, 30, 3),
    scenario: Scenario | None = None,
    model: str | ModelParams | None = None,
    cash_interest: bool = False,
    trusted_agents: bool = False,
    history_days: int = 0,
    margin_interest: bool = True,
) -> dict[str, Scorecard]
ArgumentDefaultMeaning
agentsrequiredName to agent. An agent is any object with act(obs), or a StrategySpec, which is built fresh on every call.
seedrequiredAny integer from 0 to 2**64 - 1. Fixes the market every agent meets.
universerequiredThe roster. Its order is part of the market.
macroNoneThe economy on day 0. None uses Macro()'s defaults.
days5Trading days to score.
steps_per_day6Decisions per day. The agent acts once per step.
ticks_per_step65Minutes between decisions.
cash1000000.0Starting cash per agent.
max_leverage2.0Cap on gross exposure as a multiple of net worth. None removes it.
start(9, 30, 3)Hour, minute and day of the week the first session opens on.
scenarioNoneA Scenario to drive the economy. None lets it run on its own.
modelNoneA preset name or a ModelParams. None is the default preset, pt-v20.
cash_interestFalseTrue pays uninvested cash the policy rate, a day's worth before each close.
trusted_agentsFalseTrue hands agents the live engine and portfolio in place of the read-only views. Every scorecard then says trusted.
history_days0Days to run before day 0 with nobody trading, readable through obs.history. At most 2,520. The scored days continue the warmed market.
margin_interestTrueA negative cash balance pays the policy rate before each close. False borrows for free, as 0.8.1 and earlier did, and the scorecard says free-borrowing.

Returns a dict of Scorecard, keyed by the names you passed.

evaluate builds one engine and forks it for each agent and for an untraded baseline, so every agent meets the same market and none uses liquidity another expected. Agent objects are not copied, so state an agent keeps persists across its own steps. An agent's fills in a step reach its market once, on the first tick of the next step. A bad entry in an agent's dict is refused with a line in errors and the rest trades. A return other than a dict, or an exception from act, trades nothing that step.

evaluate warns, without changing anything it runs, when an agent failed on every step, when every order in a step was a fraction of a share (weights sent as shares), and when a spec's top_k is more than half the roster.

reference_agents

The five baseline agents.

tf.reference_agents(*, seed: int = 0) -> dict[str, Agent]

seed seeds the random agent only. The keys are buy_and_hold, random, momentum, mean_reversion and oracle. The Oracle declares privileged = True and reads the model's fair value, holding the same gross exposure as the others. On pt-v20 it is a reference agent and not a ceiling, so compare agents with buy-and-hold there.

baselines.Balanced

A fixed-weight portfolio of the equities and the simulated rate indices, 60/40 by default, with a drift band. It is one of the baselines reference_agents leaves out.

tf.baselines.Balanced(
    *, equity: float = 0.6,
    bonds: dict[str, float] | None = None,
    band: float | None = 0.05,
    max_participation: float = 0.05,
)
ArgumentDefaultMeaning
equity0.6Share of net worth held across every equity in the roster, equally weighted.
bondsNoneRate index to share of net worth. None is UST2Y 0.10, UST10Y 0.20 and IGCORP 0.10.
band0.05How far a weight may drift before the agent trades every holding back to target, checked at each day's first step. None never rebalances.
max_participation0.05The largest trade in one name as a share of its average daily volume.

It needs a roster with the rate indices, from Universe.random(..., bonds=True) or Universe.with_bonds(). Without them it raises ValueError at its first step, which evaluate records in errors.

u = tf.Universe.random(20, seed=101, bonds=True)
scores = tf.evaluate(
    {"sixty_forty": tf.baselines.Balanced()},
    seed=3, universe=u, days=120,
    scenario=tf.Scenario.load("curve_shock"),
    cash_interest=True)

capture_ratio

tf.versus_buy_and_hold(scores, *, reference: str = "buy_and_hold") -> dict[str, float]
tf.capture_ratio(scores, *, oracle: str = "oracle") -> dict[str, float]
tf.capture_withheld(scores, *, oracle: str = "oracle") -> str | None
tf.oracle_is_ceiling(model=None) -> bool

versus_buy_and_hold gives each agent's P&L less buy-and-hold's in the same market, in currency. It is the comparison to quote on pt-v20. It leaves out tampered agents and warns, returning {}, when buy-and-hold did not run.

capture_ratio gives each agent's P&L as a fraction of the Oracle's: 1.0 is the Oracle's P&L. On pt-v20 it returns {} and warns, because a shock there moves fair value for good and the Oracle's P&L follows the market's month. capture_withheld returns that reason, or None where a ratio is reported. oracle_is_ceiling answers for a preset name, a ModelParams or the default.

leaderboard

tf.leaderboard(scores, by: str = "pnl") -> list[Scorecard]

Sorts scorecards by one field, best first, with tampered cards last.

rank

Run the same agents on many seeds and compare them seed by seed.

tf.rank(
    make_agents: Callable[[], dict[str, Agent | StrategySpec]],
    *, seeds: Iterable[int],
    universe: Sequence[Instrument],
    macro: Macro | None = None,
    days: int = 5,
    steps_per_day: int = 6,
    ticks_per_step: int = 65,
    cash: float = 1000000.0,
    max_leverage: float | None = 2.0,
    start: tuple[int, int, int] = (9, 30, 3),
    scenario: Scenario | None = None,
    oracle: str = "oracle",
    workers: int = 1,
    model: str | ModelParams | None = None,
    trusted_agents: bool = False,
    benchmark: str = "buy_and_hold",
    history_days: int = 0,
    margin_interest: bool = True,
) -> Ranking

The arguments shared with evaluate mean the same, and the others are these:

ArgumentDefaultMeaning
make_agentsrequiredCalled once per seed, and must return new agents each time, because an agent carried across seeds brings its state with it.
seedsrequiredThe seeds to run. Each is one market.
oracle"oracle"The entrant to measure capture against, where capture is reported.
workers1Threads to spread the seeds over. 1 runs them one after another. Results are collected in seed order either way.
benchmark"buy_and_hold"The entrant each agent's excess P&L is measured against.

Ranking

MemberMeaning
recordsOne AgentRecord per agent, keyed by name.
seedsThe seeds that ran.
benchmark, benchmark_noteThe entrant excess P&L is measured against, and how it was chosen.
capture_withheldWhy no capture ratio is reported, on pt-v20. None where one is.
tamperedAgents left out because they changed the market.
oracle, reference_pnls, unmeasurableThe capture reference, its P&L per seed, and the seeds where no capture could be formed.
model_fingerprint, universe_fingerprintThe model and roster every seed ran on.
report(), table(), as_dict()The results as a printable table, as rows, and as plain data.
separation(a, b)A paired sign test between two agents.

report() names the benchmark by its label, as in vs buy_and_hold or vs flat. On pt-v20 it also says that the Oracle ran and has no row, with its P&L over the benchmark's. An agent whose adapter refused actions gets a REFUSED line with the count and the first refusal, an agent whose act() returned something other than a mapping gets an UNUSABLE line with the count and the first seed and step, and an agent that raised gets a RAISED line.

separation(a, b) returns a dict with a, b, wins_a, wins_b, ties, paired_seeds, decisive and p_value. Both agents meet the same market on each seed, so each pair differs only in the agent. decisive is true only when one agent won on every seed, and it counts wins without testing significance.

AgentRecord

One agent's results across every seed.

MemberUnitsMeaning
name, seedsThe name it ran under, and its seeds.
pnlscurrencyP&L per seed.
median_pnlcurrencyThe median of pnls.
excess_pnls, mean_excess_pnlcurrencyP&L less the benchmark's per seed, and their mean. The number to quote on pt-v20.
seeds_aheadseedsSeeds on which the agent earned more than the benchmark.
wins, win_rateseeds, fractionSeeds on which it had the highest P&L of the ranked agents, and that as a share.
pooled_capture, median_capture, capture_range, capturesratioCapture of the Oracle's P&L, pooled, per seed and its spread. None or empty on pt-v20.
errors, seeds_with_errors, first_errorHow often the agent's code raised on each seed, on how many seeds, and the first line.
rejected, max_leverageorders, multipleRefused orders and peak leverage per seed.
refusedactionsPer seed, the actions an LLM adapter turned down on its own, such as a ticker the roster does not have, while the rest of the decision traded. They are part of rejected and are not raises.
first_refusalThe first of those refusals, prefixed with its seed, or None.
unusable, first_unusablestepsPer seed, the steps whose act() returned something other than a mapping, such as a list of pairs, and the first of them. They are not counted in errors.
failed_every_step, seeds_failedPer seed, whether every step raised or was refused and nothing traded, and on how many seeds. Such a seed has no score: it reads None in excess_pnls, does not count in seeds_ahead, and cannot win.
trusted, uses_hidden_stateHow this agent saw the market.
as_dict()The record as plain data.

tca.analyse

Run the same seed with and without the agent's orders, so every fill is priced against the market where it never traded.

tf.tca.analyse(
    agent,
    *, seed: int,
    universe: Sequence[Instrument],
    macro: Macro | None = None,
    days: int = 1,
    steps_per_day: int = 6,
    ticks_per_step: int = 65,
    cash: float = 1000000.0,
    max_leverage: float | None = 2.0,
    start: tuple[int, int, int] = (9, 30, 3),
    scenario: Scenario | None = None,
    model: str | ModelParams | None = None,
    trusted_agents: bool = False,
    history_days: int = 0,
) -> Execution

The agent is any object with act(obs) returning share quantities, as on Writing an agent. tca.analyse takes market orders only and refuses a tf.Limit, because the part of a limit order that waits fills inside a session, where the untraded run has no price to compare it with. days is 1 by default, unlike evaluate. history_days runs the market that many days before day 0 with nobody trading, as in evaluate, and both runs start day 0 from the warmed market. No scenario applies during the warm-up. Execution cost explains what the comparison measures.

Execution

What the trading cost, against the run without the agent's orders.

MemberUnitsMeaning
shortfall_bps()basis pointsCost of trading against the untraded path. Positive is a cost.
shortfallcurrencyThe same in currency.
by_step(), by_ticker()The cost per decision step and per name.
partial_fills()What was asked for against what filled.
impact_bps(ticker)basis pointsOne name's price move caused by the agent's own orders.
fillsEvery fill the agent took.
actual_final, baseline_finalcurrencyNet worth at the end, traded and untraded.
actual_path, baseline_pathcurrencyNet worth per step, traded and untraded.
moved, untouched_moved()Names whose price the agent moved, and names it never traded whose price still differs from the baseline.
steps, tickers, portfolio, seedThe run's size, roster, final portfolio and seed.
history_daysdaysThe warm-up both runs had before day 0.
model_fingerprint, universe_fingerprintThe model and roster it ran on.
as_dict()The result as plain data.

Scorecard

One agent's result on one seed, as evaluate returns it, with errors to read before the P&L.

AttributeUnitsMeaning
name, seedThe name it ran under and the seed.
pnlcurrencyFinal net worth less starting cash.
return_pctpercentP&L as a percentage of starting cash: 5.2 means 5.2%.
final_net_worthcurrencyCash plus positions marked at the last close.
equity_curvecurrencyNet worth after each day's close. The last value is final_net_worth.
max_drawdown_pctpercentThe largest fall from a running peak of equity_curve, which starts at the starting cash. Measured on closes.
sharperatioMean daily return over its standard deviation, times the square root of 252, from equity_curve and the starting cash. No risk-free rate is subtracted. None with fewer than two days, once net worth reached zero, or when the returns did not vary. Over five or ten days it is very noisy.
volatility_pctpercentAnnualized volatility of the same daily returns: their standard deviation times the square root of 252. None with fewer than two days or once net worth reached zero.
exposure_curvemultipleGross exposure as a multiple of net worth after each step's session, longs and shorts both counted. Infinite once net worth is at or below zero.
time_in_marketfractionShare of steps in exposure_curve that ended holding any position.
avg_gross_exposuremultipleMean of exposure_curve over the steps where net worth was above zero. 1.0 is fully invested and 2.0 the default leverage limit.
ruinedNet worth was at or below zero at some close.
impact_bpsbasis pointsThe agent's own footprint, against the run where it did not trade.
tradesfillsFills taken. A market order, or the part of a limit order that fills on arrival, counts once however many price levels it took, and each later fill of a waiting limit order counts again.
turnovercurrencyTraded value.
max_leveragemultipleHighest gross exposure reached, as a multiple of net worth.
rejectedordersOrders refused, and actions an LLM adapter refused on its own. Each has a line in errors.
leverage_refusalsordersHow many of rejected the leverage limit refused.
partial_fillsOne line per market order the book could not fill in full.
errorsOne line per step where the agent raised, returned something that is not an order dict, or sent a refused order.
explanationsWhat the agent's explain(day) answered, day by day.
explanation_accuracyfractionShare of days on which explain named the factor that moved prices most. None without explain.
explanation_baselinefractionWhat naming the most common winner every day scored on the same days.
explanation_edgefractionAccuracy less baseline. On pt-v20 the baseline is 0.95 to 1.0, so this is the figure to compare.
history_daysdaysThe warm-up before day 0. Left out of as_dict() when 0.
margin_interestWhether borrowing paid the policy rate.
trusted, uses_hidden_state, tamperedThe agent had the live engine, read hidden state, or changed the market outside its orders. A tampered card ranks nothing.
strategy_fingerprintsha256 of the StrategySpec, or empty for a hand-written agent.
model_fingerprint, universe_fingerprintThe model and roster it ran on.
as_dict()The scorecard as plain data.

sharpe, volatility_pct, time_in_market and avg_gross_exposure are read-only properties. They and exposure_curve are left out of as_dict(), so the traded digest hashes what it did before they were added.

The repr prints the four figures after the impact, then the counts and flags that are set. Over fewer than Scorecard.SHARPE_MIN_DAYS (20) scored days it prints sharpe=n/a (short run), because the standard error of an annualized Sharpe ratio is 3.5 or more there. The property still returns the figure: buy-and-hold in One market and the baselines reads sharpe=n/a (short run), vol=9.6%, in_market=100%, exposure=1.00x after ten days on a 1.05% gain, and its sharpe is +2.78. The counts and flags include errors=1, free-borrowing, history_days=20 and explanation=0.967 vs baseline 0.983 (edge -0.016).

RunManifest

A finished run as one document: the roster, the day-0 economy, the scenario, the order log, the strategy when it is a StrategySpec, the package version, the model's coefficients and the market digest a replay has to reproduce. Recording the market shows the round trip.

tf.RunManifest.of(
    engine: Engine,
    *, seed: int,
    universe: Sequence[Instrument],
    macro: Macro | None = None,
    scenario: Scenario | None = None,
    strategy: StrategySpec | str | None = None,
    universe_source: Any = None,
    label: str = "",
    derived_from: Any = None,
    ledger: DayLedger | None = None,
    agent_access: dict[str, Any] | None = None,
) -> RunManifest
MemberMeaning
of(engine, ...)Capture a finished run. The seed and roster are passed because an engine keeps neither. strategy takes a StrategySpec, carried in full, or a reference string for a hand-written agent, recorded in gaps. derived_from takes the Checkpoint the run branched from.
to_json(), from_json(text)The manifest as JSON text, and back.
reproduce()Replay and verify. Returns the rebuilt engine.
verify_lineage(checkpoint)Check that derived_from names this checkpoint.
describe()A readable summary.
seed, label, universe, universe_source, macro, scenarioThe inputs as recorded.
strategy, strategy_referenceThe carried spec, or the reference string of one that is not carried.
order_logEvery input the engine consumed.
model, model_fingerprintThe coefficients, and a preset's name or custom- plus a digest.
fingerprintsA digest per component, and inputs over the seed and all of them.
resultThe market digest, the days run and the draws consumed.
written_byThe package version, platform, Python version, model and era digest of the build that wrote it.
derived_fromThe checkpoint this run branched from, or None.
gaps, completeWhat a reader needs from outside the manifest, and whether that list is empty.

reproduce() raises ValidationError on a mismatch and names the component that disagreed. It first compares the era digest, a small fixed calculation, with the one in the manifest, and refuses to replay on a build that does different arithmetic. It checks the market and carries no score: result holds the market's digest, the number of days and draws_consumed, and nothing else. evaluate and rank write no manifest, so to let a reader check a score, publish the agent, the call and the seeds.