Skip to the page
RELEASES/RELEASE NOTES

Release notes

A market with no agent orders in it replays exactly under every later release if the run names its preset. A run that took the default replays only under the version that shipped it, so each section below names the preset that moved. A traded run replays exactly on its own release, and one recorded before 0.8.5 matches only up to its first trade. The full history is in the library's CHANGELOG.md.

0.8.7

2026-10-03

A patch on the 0.8 long-term support line that changes text only. Every known-answer digest is 0.8.6's and pt-v20 stays the default, so every run replays exactly as it did.

The documentation inside the package is rewritten for users. Each entry in Model parameters now says what the coefficient does, its units and which presets carry which value. The gap descriptions in the published envelope, the reasons score gives for a row it can't read, and several docstrings say the same in plain terms. A few stale claims were corrected on the way, and a test now checks every value a parameter's description states against the preset table.

0.8.6

2026-10-01

A patch on the 0.8 long-term support line, with no coefficient, default or trajectory changes. Every known-answer digest is 0.8.5's and pt-v20 stays the default, so every run replays exactly as it did under 0.8.5.

tf.envelope.check() now says why it refuses a sector-concentrated roster on pt-v20. The four-mix roster run was repeated on pt-v20, on seeds 101 to 130 at 252 and 504 days. At 252 days every mix held every shape row the bands can grade. At 504 days the S&P-like and technology-heavy mixes put the correlation of volume with the size of a move at 0.6367 and 0.6332, against a ceiling of 0.63, where the balanced roster reads 0.6266. The refusal and the roster-concentration gap in How it is measured quote this run, which is in the library as measurements/roster-shapes-pt-v20.json. A question that passes preset="pt-v19" keeps the pt-v19 grant.

0.8.5

2026-10-01

Long-term support

0.8.5 starts the first long-term support line. The 0.8 line gets bug and security fixes for 24 months, and a fix ships in a 0.8.x patch only if it leaves every known-answer digest unchanged. The library's SUPPORT.md is the policy, and Support policy lists the digests.

pt-v20 as the default

A run that took the default will not replay against 0.8.1, so pass model="pt-v19" to keep that market. A run that names its preset replays exactly, and every shipped preset from pt-v1 on can still be selected. On pt-v20 the tape follows the model price, a stock's own news and the market's plain shocks move fair value for good, fear marks fair value down while the VIX is above 40, and agents trade in a book with depth. It passed all 40 long-run rows registered for it on exam seeds it had not been run on, all 19 statistics of the one-year table, and 14 of 14 at two years. How it is measured describes the grade. Nearest their edges are the two-year correlation of volume with the size of a move, at 0.627 against a ceiling of 0.63, a macro timing rule, at 92% of its tolerance, and the two-year yield, which moves 3.87 bp a day against a real 5.23.

Economic data published late

On pt-v20 the business-cycle phase reaches macro_fields["cycle"], trace rows and the LLM adapters' payload 252 sessions after it turns, as the NBER dates a turn about a year late. GDP growth is a quarterly figure released 21 sessions after the quarter ends. Prices read the true state, and code that needs it reads state_snapshot()["economy"], which a sandboxed agent cannot.

No capture ratio on pt-v20

A shock on pt-v20 moves fair value for good, so knowing fair value leaves little edge and the Oracle's P&L follows the market's month. capture_ratio returns {} and warns why, versus_buy_and_hold gives each agent's P&L minus buy-and-hold's, and rank sorts on mean_excess_pnl. pt-v19 and earlier still report capture.

Orders from act

A number is a market order for that many shares. tf.Limit(quantity, price) and tf.Cancel() now work in evaluate and rank as they did in World. In evaluate a bad entry is refused with a line in the scorecard's errors and the rest of the dict trades. True and "100" as quantities, which traded before, are now refused, and a falsy return such as [] or 0 is an error where it passed as no trade. Return None or {}. Agents and evaluation has the rules.

A read-only market for agents

obs.engine is a MarketView and obs.portfolio a PortfolioView in evaluate, rank, World and tca.analyse. Anything else raises tf.SandboxError. An agent that changes the market or copies the engine is scored tampered, and rank leaves it out. privileged = True gives an agent obs.hidden, and trusted_agents=True hands every agent the live engine, and the scorecard marks both. The check catches writes and copies only. An agent can still read the seed and roster from the harness's frames, build a second engine and run it ahead, so run code you did not write in a separate process.

Warm-up history

evaluate, rank and World take history_days=N, which runs the market N days with nobody trading before the scored days. obs.history holds a daily bar per name and the published macro figures for every closed day, so a 20-day breakout rule can trade on day 1 instead of day 21. With history_days=0, the default, nothing changes. tca.analyse takes history_days too.

Margin interest

A negative cash balance pays the policy rate before each close in evaluate, rank and World, so a levered agent scores less than it did under 0.8.1. margin_interest=False borrows for free and the scorecard says free-borrowing.

Explanation accuracy and its baseline

The scorer asks which of ten factors moved prices most each day, reads the attribution after the close, and no longer counts fair_value_shift, which moves no price. On pt-v20 random_noise wins almost every day, so naming it every day scores 0.95 to 1.0. The scorecard carries explanation_baseline and explanation_edge and prints all three. Quote the edge.

New scorecard fields

The scorecard gains equity_curve, max_drawdown_pct, ruined, leverage_refusals, partial_fills, history_days, margin_interest and exposure_curve, the gross exposure after each step. Four read-only properties come from those curves: sharpe and volatility_pct, annualized from the daily returns with no risk-free rate subtracted, and time_in_market and avg_gross_exposure. The repr prints all four and counts errors. Agents and evaluation defines every field with its units.

Refusals in rank

An action an LLM adapter refuses on its own, such as a ticker the roster does not have, counts in Scorecard.rejected. AgentRecord.refused holds the count per seed, and the report gives the agent a REFUSED line with the first refusal. A seed on which an agent failed at every step has no score, so it no longer counts as ahead of a falling market. The report names the benchmark it read, and on pt-v20 it says that the Oracle ran and has no row.

Bar volume

Engine.bars() summed a running total, so a day bar read about two hundred times the day's volume. A bar's volume is now the shares traded in it, at every grain, so a day's tick rows add up to its five-minute bars and its day bar. If you read the tick column as a running total, take its cumulative sum per name and day. Prices and digests are unchanged. On pt-v20 volume_abs_return_corr moves from 0.508 to 0.596 at one year and from 0.561 to 0.627 at two, and volume_change_acf1 from -0.254 to -0.268 at one year, all still in their bands.

LLM decisions and payload

Decision schema 2 lets an action carry a limit_price, which becomes a tf.Limit, and side: "CANCEL", which becomes a tf.Cancel(). A bad action is refused on its own and the rest of the decision trades. Observation payload 1 is frozen for the 0.8.x line. portfolio.gross_exposure is renamed leverage, portfolio.open_orders lists waiting limit orders, and return_5d covers five full days. A recording carries both schema versions and a replay refuses one made under another, so recordings made before 0.8.5 do not replay. LLM adapters and MCP has the contract.

Fingerprint battery version 2

Seven 120-day markets, one per shipped scenario including curve_shock, so a day-50 shock has 70 days after it. tf.battery(1) still builds the old six markets, and a fingerprint compares only with one taken on the same version.

A digest for a traded run

The known-answer script now hashes one run through tf.evaluate on pt-v20: the five reference agents and an agent that sends and cancels limit orders, with every order, fill and scorecard field. It is checked on all five platforms beside the other digests.

Fills reach the market once

Every harness passed an agent's fills to run_session as order_flow, which the session held on every tick, so one order counted 65 times and agents were marked to their own impact. They now arrive once, as fills. On pt-v19 the spec mean-reversion rule had beaten buy-and-hold on all 20 suite markets by a median 42 points in 60 days, and with fills applied once it reads +0.5 points, ahead in 10 of 20. Every traded result moves. run_session(order_flow=...) now raises, so pass fills=, or flow_per_tick= for a standing rate. Logs, checkpoints and manifests from 0.8.x replay as they ran. Engine and data has the change.

A book with depth

Seven ModelParams dials give the book depth priced by size, consumption and refill, resting limit orders and a permanent impact per agent. All seven are 0 on every earlier preset, and pt-v20 turns them on. Engine and data says what each does to an order. A resting order fills against the model's flow only at a price inside the maker's quote for that tick. A resting order that the maker's re-quote crosses trades at the maker's price and is recorded as liquidity="taker". tf.Limit and tf.Cancel compare by value, so agree no longer reports two forks that sent the same limit orders as different.

Rate indices

Universe.random(40, seed=1, bonds=True) adds UST2Y, UST10Y and IGCORP, priced off the engine's own curve. None is a real security, and every equity price is the same with or without them. The curve_shock scenario moves the whole curve 200 bp in one day, baselines.Balanced holds a 60/40 book, and cash_interest=True pays idle cash the policy rate. Engine and data has the pricing.

64-bit seeds

Seeds take any 64-bit integer, and every seed below 2**32 gives the market it gave before. Draw a sealed seed with secrets.randbits(64).

Recalibrated crisis scenarios

On pt-v20 recession takes the index down 44.7% at 120 sessions (2008 fell 45%) and now ends, with the cycle at trough on day 365. liquidity_crisis falls 33.9% at worst, as March 2020 did, though more slowly. Scenarios says why.

Gym resets

The gym environment draws a new market on each reset(), where a reset without a seed used to replay the constructor's market. info["seed"] names the episode's seed. An action over the leverage cap is scaled down to 1.96x gross, with info["scaled"] set.

Speed

evaluate of five momentum strategies on 40 names over 20 days fell from 11.0 to 3.95 seconds of CPU, and building a pt-v20 engine from 0.70 to about 0.01 seconds.

Python 3.12 and 3.13

3.12 changed sum() over floats, so an agent's orders could split from 3.11's in the last digit. tradefloor now adds floats left to right on every version, so 3.12 and 3.13 match 3.11.

Short-lag clustering

Volatility clustering reads about a quarter of real at lag 1, so the decay-shape gap now covers abs_return_acf1 and abs_return_acf5 as well as abs_return_acf20, and tf.envelope.check() refuses a question that relies on any of the three. How it is measured lists the gaps.

Examples and the MCP server

examples/08-claude-agent.py replays a committed Claude run by default, so it runs with no key and no network, and TRADEFLOOR_LIVE_EXAMPLES=1 with a key calls Claude once per simulated day. The MCP server refuses a stress run that ends before its scenario's first event, where it returned a difference of 0.0, and refuses strategies named after a baseline. list_scenarios gives each scenario's first and last event day, the constructors are rate_ramp and vix_shock, and pip install "tradefloor[mcp]" now installs pyarrow, which explain needs.

The model written down

The model specification states pt-v20 as equations, each with its source line, and gives every coefficient its value and how it was set. STATISTICS.md names the statistic sets behind every count the site quotes.

Rust crate API changes

The crate takes the package's version, so Cargo treats 0.8.5 as a compatible update to 0.8.1, and it is not. Seeds are u64, several public structs have new fields and are #[non_exhaustive], and Engine::new builds a pt-v20 market. Pin tradefloor = "=0.8.1" to stay on the old API. The CHANGELOG lists every changed signature.

Smaller fixes

An agent no longer trades with itself in the settlement book. A resting fill can sit outside the day's high and low, because a bar keeps only each tick's last print. The Scorecard repr prints sharpe=n/a (short run) for a run of fewer than 20 scored days. An act() that returns something other than a mapping is reported on an UNUSABLE line and held in AgentRecord.unusable. The scripts and output that graded pt-v20 are in the library's validation/pt-v20/.

Known gaps

The worst month of the 2020 replay is about 19% milder than the real one. With the VIX held at 65 the market is 5.1 times as volatile as with it held at 5, against 6.2 times in real markets.

0.8.1

2026-09-23

Text only, with no coefficient, default or trajectory changes. The README, the messages tf.envelope.check() prints, the scenario target notes and the parameter descriptions now describe pt-v19, where some still quoted older presets. Re-measured on pt-v19, the decay slope of volatility clustering reads -0.515 against a real -0.436, and the memory holds to lag 20. A driven 2020-21 scenario moves prices at about a fifth of the real size, and qe_pe_boost moves nothing on pt-v16 and later.

0.8.0

2026-09-23

pt-v19 became the default. A run that took the default will not replay against 0.7.x, so pass model="pt-v18" to keep the 0.7.x market. pt-v19 meets all fifteen criteria of a long-run check (thirty 21-year histories, and 2008 and 2020 replayed with the real VIX forced in), where pt-v18 meets eight. tf.preset_record("pt-v19")["long_run"] holds every row. Over 21 years it has 1.35 bear markets a decade against a real 1.12, index volatility of 16.6% against 18.1%, and the VIX above 30 on 8.1% of sessions against 8.2%. The 2008 replay falls 41% against the real 57%.

ModelParams.from_preset refuses an override that breaks one of pt-v19's three identities, such as garch_alpha=0.07 alone, because garch_beta follows from it. ModelParams.from_preset_unchecked builds it anyway, and ModelParams.identity_breaks lists what broke.

A recorded transcript names its preset under meta["model_preset"], and replaying it against a different preset raises ReplayMiss. A manifest or checkpoint written under 0.7.x is refused, because its probe simulation runs the default preset. state_hash covers thirteen more fields, and a custom parameter set's fingerprint changes between releases.

Realism is graded on the ruled bands, each read from the longest real record for its row. envelope.certified(), envelope.score() and facts.report() take basis="shipped" for the old table. New calls: Engine.session_news(), Engine.session_tick, Scenario(vix_sets_variance=True) and hold(epicentre=...).

0.7.1

2026-09-08

Documentation only. No behavior changes, and every published digest holds.

0.7.0

2026-09-08

pt-v18 became the default, and the first default to hold every certified row: the index returns +5.80 percent a year where pt-v16 lost 13.64. The steady-state crisis lever reads 6.53x against real markets' 6.16x. tf.preset_names() lists the presets the engine resolves, and tf.preset_record reads pt-v1 through pt-v18 from JSON in the wheel.

Every Universe.random roster re-rolls, so pin 0.6.2 for the old draws. An EDGAR roster and a roster you build yourself stay as they were. A supplied macro_state now survives construction, the first central-bank meeting and OPEC decision fall inside a 252-day year, and to_instruments reads the discount rate from the model. Engine.explain decomposes a name's day down to the draws that seeded it.

0.6.2

2026-09-01

A scenario's liquidity shock now reaches the agent's volume and order cap. World(on_refusal="skip") records an unusable response and carries on. Every adapter takes prior=, a recording consulted before the provider. resample() measures an agent's own noise floor. fetch(ciks=) returns exactly those EDGAR filers, and saved files are the same bytes on Windows.

0.6.1

2026-08-31

tradefloor.integrations gains adapters for the OpenAI Agents SDK, PydanticAI and LangGraph, and one for any plain Python function. A Transcript keys each exchange by a digest of the exact input, so a recorded run replays with no framework installed and no network. The examples in examples/integrations/ run offline. LLM adapters and MCP has the contract.

0.6.0

2026-08-30

pt-v16 became the default. It couples the VIX to the market's own realized volatility. tf.Scenario and six packaged scenarios arrived, with the tradefloor scenario command. tradefloor.counterfactual runs one agent in two worlds that differ by one variable, with World, agree() and compare().

0.5.0

2026-08-28

The library was renamed from pretium, which published through 0.4.3 and stays on PyPI and crates.io. A result computed under one of those versions is cited as pretium at that exact version. The rename changes no behavior, and preset names stay pt-v1 through pt-v15. pt-v15 is selectable by name, and pt-v14 remains the default.

0.4.3

2026-08-28

Restores pt-v13 and pt-v14 exactly as they were in 0.4.0 and 0.4.1, after 0.4.2 moved their dollar safe-haven gate. The gate is its own dial now, usd_crisis_vix_threshold.

0.4.2

2026-08-28

Three fixes with no trajectory change: crisis_vix_threshold now gates the dollar's safe-haven drift as well as gold, DayAdvanceOutcome carries a meeting's decision, and daily_credit_floor_gain ships at 0.0.

0.4.1

2026-08-28

pt-v13 and pt-v14 now report the 60-day mispricing half-life they run, where tf.model_preset() said 68.26. No trajectory moves.

0.4.0

2026-08-28

pt-v14 became the default. Over 13 seed blocks it holds the two-year panel fully in band on 11, where pt-v12 held 3, and puts 137 of 138 roster shapes in band.

0.3.0

2026-08-27

pt-v12 became the default, the first preset in band on all 14 statistics over two years as well as one. Results recorded without naming a preset differ from here on, and model="pt-v10" gives the old market.