Agents and evaluation
Reference for the agent protocol, tf.evaluate, tf.rank, the baselines, tf.tca.analyse and the records they return: Scorecard, Ranking, AgentRecord, Execution and RunManifest. Worked examples are in the guides: Write an agent builds an agent and sizes its orders, Compare strategies ranks agents across seeds and prices their trading, and Record and replay publishes a run a reader can check.
Writing an agent
An agent is any Python object with an act(obs) method. At each decision step the harness passes it an observation and reads back orders. By default a run has 6 decision steps a day, 65 simulated minutes apart, and each agent starts with 1,000,000 in cash and may hold positions worth up to twice its net worth. obs.step counts steps across the whole run, and obs.step_of_day, obs.is_first_step_of_day and obs.is_last_step_of_day count within the day. An agent may also declare privileged = True and keep an explain(day) method, both described below. A first agent builds and scores one on a seeded market.
What act returns
A dict of ticker to order, where each value is one of these:
| Value | Order |
|---|---|
1000 or -500 | A market order for that many shares. It fills against the book at once, and what the book cannot fill is listed in partial_fills. |
tf.Limit(500, 101.25) | Buy up to 500 shares at 101.25 or better, or sell with a negative quantity. What does not fill waits in the book, behind the shares already at that price, until it fills, you cancel it, or a new limit on the ticker replaces it. |
tf.Cancel() | Cancel every order of yours still waiting on that ticker. |
Return None or {} to trade nothing. Values are shares, never weights, and any finite number of shares is accepted, fractions included. To hold a target weight, send the target holding less the shares already held and the shares in orders still waiting on the ticker, with the target rounded to whole shares if you want them:
target = int(weight * obs.portfolio.net_worth() / obs.price(ticker)) order = target - obs.position(ticker) - waiting
The target alone, such as 0.2 * net_worth / price, is the holding to reach, and sent as an order while a position exists it buys the position again. Sizing to a target weight works through an example with a waiting order.
There are no stop or bracket orders, so check a stop yourself at each step. A refused entry, such as a string or NaN, gets a line in the scorecard's errors and the rest of the dict still trades, so read errors before the P&L.
What the agent sees
The observation holds the roster and last prices, each name's order book and average volume, the agent's own positions, cash and waiting orders, daily bars and the published economic figures in obs.history, and a view of the market in obs.engine. What obs.engine and obs.portfolio serve depends on the agent's access: ordinary by default, privileged with privileged = True on the agent, or trusted with trusted_agents=True on the call. The levels are the same under evaluate, rank, World and tca.analyse, and The agent's view sets out what each one serves.
Ordinary and privileged agents get the read-only MarketView and PortfolioView, which raise tf.SandboxError for anything they do not serve, recorded in errors. obs.hidden is None unless the agent is privileged, when it is a read-only HiddenState that also carries the size of each day's news. obs.engine.bars() refuses under evaluate, which records no days, so daily bars come from obs.history. An LLM framework run through an adapter sees only the payload the adapter builds, the same at every level: What the model sees.
Warm-up history
history_days=N on evaluate, rank, World or tca.analyse runs the market for N days before day 0 with nobody trading, and obs.history holds those days at the first decision, labeled -N to -1. Each scored day is added after its last step. N is at most 2,520, ten years of 252 days, and no scenario applies during the warm-up. Warm-up history runs a breakout rule on it.
Against buy-and-hold, across seeds
evaluate scores one market, which is one sample. Compare each agent with buy-and-hold on the same market with versus_buy_and_hold, and run rank over many seeds before calling a winner. Compare strategies does both and reads the paired sign test. To put several agents in one market, use a World.
Explaining moves
An agent may also have explain(day). The harness calls it after each close and expects the name of the factor that moved prices most that day. On pt-v20 random_noise wins almost every day, so the scorecard reports explanation_baseline, what naming the most common winner every day scored, and explanation_edge, the accuracy less that baseline, which is the figure to quote.
Execution cost
tf.tca.analyse prices every fill against the same seed run without the agent's orders. Execution cost in Compare strategies prices a round trip and says what the figure measures, and tca.analyse below has the arguments.
evaluate
Score each agent on its own copy of the same market.
tf.evaluate(
agents: dict[str, Agent | StrategySpec],
*, seed: int,
universe: Sequence[Instrument],
macro: Macro | None = None,
days: int = 5,
steps_per_day: int = 6,
ticks_per_step: int = 65,
cash: float = 1000000.0,
max_leverage: float | None = 2.0,
start: tuple[int, int, int] = (9, 30, 3),
scenario: Scenario | None = None,
model: str | ModelParams | None = None,
cash_interest: bool = False,
trusted_agents: bool = False,
history_days: int = 0,
margin_interest: bool = True,
) -> dict[str, Scorecard]| Argument | Default | Meaning |
|---|---|---|
agents | required | Name to agent. An agent is any object with act(obs), or a StrategySpec, which is built fresh on every call. |
seed | required | Any integer from 0 to 2**64 - 1. Fixes the market every agent meets. |
universe | required | The roster. Its order is part of the market. |
macro | None | The economy on day 0. None uses Macro()'s defaults. |
days | 5 | Trading days to score. |
steps_per_day | 6 | Decisions per day. The agent acts once per step. |
ticks_per_step | 65 | Minutes between decisions. |
cash | 1000000.0 | Starting cash per agent. |
max_leverage | 2.0 | Cap on gross exposure as a multiple of net worth. None removes it. |
start | (9, 30, 3) | Hour, minute and day of the week the first session opens on. |
scenario | None | A Scenario to drive the economy. None lets it run on its own. |
model | None | A preset name or a ModelParams. None is the default preset, pt-v20. |
cash_interest | False | True pays uninvested cash the policy rate, a day's worth before each close. |
trusted_agents | False | True hands agents the live engine and portfolio in place of the read-only views. Every scorecard then says trusted. |
history_days | 0 | Days to run before day 0 with nobody trading, readable through obs.history. At most 2,520. The scored days continue the warmed market. |
margin_interest | True | A negative cash balance pays the policy rate before each close. False borrows for free, as 0.8.1 and earlier did, and the scorecard says free-borrowing. |
Returns a dict of Scorecard, keyed by the names you passed.
evaluate builds one engine and forks it for each agent and for an untraded baseline, so every agent meets the same market and none uses liquidity another expected. Agent objects are not copied, so state an agent keeps persists across its own steps. An agent's fills in a step reach its market once, on the first tick of the next step. A bad entry in an agent's dict is refused with a line in errors and the rest trades. A return other than a dict, or an exception from act, trades nothing that step.
evaluate warns, without changing anything it runs, when an agent failed on every step, when every order in a step was a fraction of a share (weights sent as shares), and when a spec's top_k is more than half the roster.
reference_agents
The five baseline agents.
tf.reference_agents(*, seed: int = 0) -> dict[str, Agent]
seed seeds the random agent only. The keys are buy_and_hold, random, momentum, mean_reversion and oracle. The Oracle declares privileged = True and reads the model's fair value, holding the same gross exposure as the others. On pt-v20 it is a reference agent and not a ceiling, so compare agents with buy-and-hold there.
baselines.Balanced
A fixed-weight portfolio of the equities and the simulated rate indices, 60/40 by default, with a drift band. It is one of the baselines reference_agents leaves out.
tf.baselines.Balanced(
*, equity: float = 0.6,
bonds: dict[str, float] | None = None,
band: float | None = 0.05,
max_participation: float = 0.05,
)| Argument | Default | Meaning |
|---|---|---|
equity | 0.6 | Share of net worth held across every equity in the roster, equally weighted. |
bonds | None | Rate index to share of net worth. None is UST2Y 0.10, UST10Y 0.20 and IGCORP 0.10. |
band | 0.05 | How far a weight may drift before the agent trades every holding back to target, checked at each day's first step. None never rebalances. |
max_participation | 0.05 | The largest trade in one name as a share of its average daily volume. |
It needs a roster with the rate indices, from Universe.random(..., bonds=True) or Universe.with_bonds(). Without them it raises ValueError at its first step, which evaluate records in errors.
u = tf.Universe.random(20, seed=101, bonds=True)
scores = tf.evaluate(
{"sixty_forty": tf.baselines.Balanced()},
seed=3, universe=u, days=120,
scenario=tf.Scenario.load("curve_shock"),
cash_interest=True)capture_ratio
tf.versus_buy_and_hold(scores, *, reference: str = "buy_and_hold") -> dict[str, float] tf.capture_ratio(scores, *, oracle: str = "oracle") -> dict[str, float] tf.capture_withheld(scores, *, oracle: str = "oracle") -> str | None tf.oracle_is_ceiling(model=None) -> bool
versus_buy_and_hold gives each agent's P&L less buy-and-hold's in the same market, in currency. It is the comparison to quote on pt-v20. It leaves out tampered agents and warns, returning {}, when buy-and-hold did not run.
capture_ratio gives each agent's P&L as a fraction of the Oracle's: 1.0 is the Oracle's P&L. On pt-v20 it returns {} and warns, because a shock there moves fair value for good and the Oracle's P&L follows the market's month. capture_withheld returns that reason, or None where a ratio is reported. oracle_is_ceiling answers for a preset name, a ModelParams or the default.
leaderboard
tf.leaderboard(scores, by: str = "pnl") -> list[Scorecard]
Sorts scorecards by one field, best first, with tampered cards last.
rank
Run the same agents on many seeds and compare them seed by seed.
tf.rank(
make_agents: Callable[[], dict[str, Agent | StrategySpec]],
*, seeds: Iterable[int],
universe: Sequence[Instrument],
macro: Macro | None = None,
days: int = 5,
steps_per_day: int = 6,
ticks_per_step: int = 65,
cash: float = 1000000.0,
max_leverage: float | None = 2.0,
start: tuple[int, int, int] = (9, 30, 3),
scenario: Scenario | None = None,
oracle: str = "oracle",
workers: int = 1,
model: str | ModelParams | None = None,
trusted_agents: bool = False,
benchmark: str = "buy_and_hold",
history_days: int = 0,
margin_interest: bool = True,
) -> RankingThe arguments shared with evaluate mean the same, and the others are these:
| Argument | Default | Meaning |
|---|---|---|
make_agents | required | Called once per seed, and must return new agents each time, because an agent carried across seeds brings its state with it. |
seeds | required | The seeds to run. Each is one market. |
oracle | "oracle" | The entrant to measure capture against, where capture is reported. |
workers | 1 | Threads to spread the seeds over. 1 runs them one after another. Results are collected in seed order either way. |
benchmark | "buy_and_hold" | The entrant each agent's excess P&L is measured against. |
Ranking
| Member | Meaning |
|---|---|
records | One AgentRecord per agent, keyed by name. |
seeds | The seeds that ran. |
benchmark, benchmark_note | The entrant excess P&L is measured against, and how it was chosen. |
capture_withheld | Why no capture ratio is reported, on pt-v20. None where one is. |
tampered | Agents left out because they changed the market. |
oracle, reference_pnls, unmeasurable | The capture reference, its P&L per seed, and the seeds where no capture could be formed. |
model_fingerprint, universe_fingerprint | The model and roster every seed ran on. |
report(), table(), as_dict() | The results as a printable table, as rows, and as plain data. |
separation(a, b) | A paired sign test between two agents. |
report() names the benchmark by its label, as in vs buy_and_hold or vs flat. On pt-v20 it also says that the Oracle ran and has no row, with its P&L over the benchmark's. An agent whose adapter refused actions gets a REFUSED line with the count and the first refusal, an agent whose act() returned something other than a mapping gets an UNUSABLE line with the count and the first seed and step, and an agent that raised gets a RAISED line.
separation(a, b) returns a dict with a, b, wins_a, wins_b, ties, paired_seeds, decisive and p_value. Both agents meet the same market on each seed, so each pair differs only in the agent. decisive is true only when one agent won on every seed, and it counts wins without testing significance.
AgentRecord
One agent's results across every seed.
| Member | Units | Meaning |
|---|---|---|
name, seeds | The name it ran under, and its seeds. | |
pnls | currency | P&L per seed. |
median_pnl | currency | The median of pnls. |
excess_pnls, mean_excess_pnl | currency | P&L less the benchmark's per seed, and their mean. The number to quote on pt-v20. |
seeds_ahead | seeds | Seeds on which the agent earned more than the benchmark. |
wins, win_rate | seeds, fraction | Seeds on which it had the highest P&L of the ranked agents, and that as a share. |
pooled_capture, median_capture, capture_range, captures | ratio | Capture of the Oracle's P&L, pooled, per seed and its spread. None or empty on pt-v20. |
errors, seeds_with_errors, first_error | How often the agent's code raised on each seed, on how many seeds, and the first line. | |
rejected, max_leverage | orders, multiple | Refused orders and peak leverage per seed. |
refused | actions | Per seed, the actions an LLM adapter turned down on its own, such as a ticker the roster does not have, while the rest of the decision traded. They are part of rejected and are not raises. |
first_refusal | The first of those refusals, prefixed with its seed, or None. | |
unusable, first_unusable | steps | Per seed, the steps whose act() returned something other than a mapping, such as a list of pairs, and the first of them. They are not counted in errors. |
failed_every_step, seeds_failed | Per seed, whether every step raised or was refused and nothing traded, and on how many seeds. Such a seed has no score: it reads None in excess_pnls, does not count in seeds_ahead, and cannot win. | |
trusted, uses_hidden_state | How this agent saw the market. | |
as_dict() | The record as plain data. |
tca.analyse
Run the same seed with and without the agent's orders, so every fill is priced against the market where it never traded.
tf.tca.analyse(
agent,
*, seed: int,
universe: Sequence[Instrument],
macro: Macro | None = None,
days: int = 1,
steps_per_day: int = 6,
ticks_per_step: int = 65,
cash: float = 1000000.0,
max_leverage: float | None = 2.0,
start: tuple[int, int, int] = (9, 30, 3),
scenario: Scenario | None = None,
model: str | ModelParams | None = None,
trusted_agents: bool = False,
history_days: int = 0,
) -> ExecutionThe agent is any object with act(obs) returning share quantities, as on Writing an agent. tca.analyse takes market orders only and refuses a tf.Limit, because the part of a limit order that waits fills inside a session, where the untraded run has no price to compare it with. days is 1 by default, unlike evaluate. history_days runs the market that many days before day 0 with nobody trading, as in evaluate, and both runs start day 0 from the warmed market. No scenario applies during the warm-up. Execution cost explains what the comparison measures.
Execution
What the trading cost, against the run without the agent's orders.
| Member | Units | Meaning |
|---|---|---|
shortfall_bps() | basis points | Cost of trading against the untraded path. Positive is a cost. |
shortfall | currency | The same in currency. |
by_step(), by_ticker() | The cost per decision step and per name. | |
partial_fills() | What was asked for against what filled. | |
impact_bps(ticker) | basis points | One name's price move caused by the agent's own orders. |
fills | Every fill the agent took. | |
actual_final, baseline_final | currency | Net worth at the end, traded and untraded. |
actual_path, baseline_path | currency | Net worth per step, traded and untraded. |
moved, untouched_moved() | Names whose price the agent moved, and names it never traded whose price still differs from the baseline. | |
steps, tickers, portfolio, seed | The run's size, roster, final portfolio and seed. | |
history_days | days | The warm-up both runs had before day 0. |
model_fingerprint, universe_fingerprint | The model and roster it ran on. | |
as_dict() | The result as plain data. |
Scorecard
One agent's result on one seed, as evaluate returns it, with errors to read before the P&L.
| Attribute | Units | Meaning |
|---|---|---|
name, seed | The name it ran under and the seed. | |
pnl | currency | Final net worth less starting cash. |
return_pct | percent | P&L as a percentage of starting cash: 5.2 means 5.2%. |
final_net_worth | currency | Cash plus positions marked at the last close. |
equity_curve | currency | Net worth after each day's close. The last value is final_net_worth. |
max_drawdown_pct | percent | The largest fall from a running peak of equity_curve, which starts at the starting cash. Measured on closes. |
sharpe | ratio | Mean daily return over its standard deviation, times the square root of 252, from equity_curve and the starting cash. No risk-free rate is subtracted. None with fewer than two days, once net worth reached zero, or when the returns did not vary. Over five or ten days it is very noisy. |
volatility_pct | percent | Annualized volatility of the same daily returns: their standard deviation times the square root of 252. None with fewer than two days or once net worth reached zero. |
exposure_curve | multiple | Gross exposure as a multiple of net worth after each step's session, longs and shorts both counted. Infinite once net worth is at or below zero. |
time_in_market | fraction | Share of steps in exposure_curve that ended holding any position. |
avg_gross_exposure | multiple | Mean of exposure_curve over the steps where net worth was above zero. 1.0 is fully invested and 2.0 the default leverage limit. |
ruined | Net worth was at or below zero at some close. | |
impact_bps | basis points | The agent's own footprint, against the run where it did not trade. |
trades | fills | Fills taken. A market order, or the part of a limit order that fills on arrival, counts once however many price levels it took, and each later fill of a waiting limit order counts again. |
turnover | currency | Traded value. |
max_leverage | multiple | Highest gross exposure reached, as a multiple of net worth. |
rejected | orders | Orders refused, and actions an LLM adapter refused on its own. Each has a line in errors. |
leverage_refusals | orders | How many of rejected the leverage limit refused. |
partial_fills | One line per market order the book could not fill in full. | |
errors | One line per step where the agent raised, returned something that is not an order dict, or sent a refused order. | |
explanations | What the agent's explain(day) answered, day by day. | |
explanation_accuracy | fraction | Share of days on which explain named the factor that moved prices most. None without explain. |
explanation_baseline | fraction | What naming the most common winner every day scored on the same days. |
explanation_edge | fraction | Accuracy less baseline. On pt-v20 the baseline is 0.95 to 1.0, so this is the figure to compare. |
history_days | days | The warm-up before day 0. Left out of as_dict() when 0. |
margin_interest | Whether borrowing paid the policy rate. | |
trusted, uses_hidden_state, tampered | The agent had the live engine, read hidden state, or changed the market outside its orders. A tampered card ranks nothing. | |
strategy_fingerprint | sha256 of the StrategySpec, or empty for a hand-written agent. | |
model_fingerprint, universe_fingerprint | The model and roster it ran on. | |
as_dict() | The scorecard as plain data. |
sharpe, volatility_pct, time_in_market and avg_gross_exposure are read-only properties. They and exposure_curve are left out of as_dict(), so the traded digest hashes what it did before they were added.
The repr prints the four figures after the impact, then the counts and flags that are set. Over fewer than Scorecard.SHARPE_MIN_DAYS (20) scored days it prints sharpe=n/a (short run), because the standard error of an annualized Sharpe ratio is 3.5 or more there. The property still returns the figure: buy-and-hold in One market and the baselines reads sharpe=n/a (short run), vol=9.6%, in_market=100%, exposure=1.00x after ten days on a 1.05% gain, and its sharpe is +2.78. The counts and flags include errors=1, free-borrowing, history_days=20 and explanation=0.967 vs baseline 0.983 (edge -0.016).
RunManifest
A finished run as one document: the roster, the day-0 economy, the scenario, the order log, the strategy when it is a StrategySpec, the package version, the model's coefficients and the market digest a replay has to reproduce. Recording the market shows the round trip.
tf.RunManifest.of(
engine: Engine,
*, seed: int,
universe: Sequence[Instrument],
macro: Macro | None = None,
scenario: Scenario | None = None,
strategy: StrategySpec | str | None = None,
universe_source: Any = None,
label: str = "",
derived_from: Any = None,
ledger: DayLedger | None = None,
agent_access: dict[str, Any] | None = None,
) -> RunManifest| Member | Meaning |
|---|---|
of(engine, ...) | Capture a finished run. The seed and roster are passed because an engine keeps neither. strategy takes a StrategySpec, carried in full, or a reference string for a hand-written agent, recorded in gaps. derived_from takes the Checkpoint the run branched from. |
to_json(), from_json(text) | The manifest as JSON text, and back. |
reproduce() | Replay and verify. Returns the rebuilt engine. |
verify_lineage(checkpoint) | Check that derived_from names this checkpoint. |
describe() | A readable summary. |
seed, label, universe, universe_source, macro, scenario | The inputs as recorded. |
strategy, strategy_reference | The carried spec, or the reference string of one that is not carried. |
order_log | Every input the engine consumed. |
model, model_fingerprint | The coefficients, and a preset's name or custom- plus a digest. |
fingerprints | A digest per component, and inputs over the seed and all of them. |
result | The market digest, the days run and the draws consumed. |
written_by | The package version, platform, Python version, model and era digest of the build that wrote it. |
derived_from | The checkpoint this run branched from, or None. |
gaps, complete | What a reader needs from outside the manifest, and whether that list is empty. |
reproduce() raises ValidationError on a mismatch and names the component that disagreed. It first compares the era digest, a small fixed calculation, with the one in the manifest, and refuses to replay on a build that does different arithmetic. It checks the market and carries no score: result holds the market's digest, the number of days and draws_consumed, and nothing else. evaluate and rank write no manifest, so to let a reader check a score, publish the agent, the call and the seeds.