Skip to the page
MODEL AND VALIDATION/HOW IT IS MEASURED

How it is measured

Every check and its pass mark is written down and committed before the grade, the grade runs on seeds the tuning never used, and the scripts and their output are published. pt-v20 passed all 40 of its registered rows that way, and all 19 statistics of the one-year table are inside the ranges real markets show.

Registration, then one grade

1Measure real marketsThe S&P 500 since 1928, the VIX since 1990, 40 large US stocks, Treasury and corporate yields, Federal Reserve decisions and Shiller's long-run earnings give the figures a trader would notice, such as how many days a decade the index falls more than 5%. Each becomes a check with a pass mark wide enough to allow for how much the real figure moves from one decade to the next.
2Write the rulesMost behaviors come from a published source, such as the square-root cost of size. A new rule goes in switched off, and a test that runs one fixed market on every shipped preset proves that every other preset's output did not change by a single bit.
3Screen on practice seedsCandidate settings run in grids over many simulated histories. Every market comes from a seed, and the seeds are split: practice seeds are used while choosing, and exam seeds are kept back for the grade.
4Try to cheat itSome checks are strategies built to find leaks: trading on a headline 5 minutes late, rules that read only past prices, value and momentum screens, timing the market on published economic data, and trading on a rate decision. Each passes only if it earns about nothing, as it would in a real market. An independent audit found two leaks, and both were fixed in the model.
5RegisterEvery check, its pass mark and the exact settings to be graded are committed. Nothing on the list can change after the results come in.
6Grade onceThe model runs once on the exam seeds and has to pass every row. If it fails one it does not ship, and a change needs a new registration and fresh seeds before another grade.
7FreezeA model that passes ships as a preset with a fingerprint of its coefficients and known-answer digests that every release checks on five platforms. A shipped preset never changes.

The seeds

For pt-v20 the final settings were chosen on practice seeds 201 to 230, 501 to 530 and 801 to 830. The grade ran on exam seeds 101 to 130, 401 to 430 and 701 to 730: 90 histories of 21 years with nothing imposed, plus replays of 2008 and 2020 with the real VIX, the real 2020-21 and 2022 economies driven through the model, the packaged recession, and the leak-finding strategies.

The 40 registered rows

Each row is something a user would notice, with a tolerance that is easy to read: within 30% on the size of a crash, half to twice the real rate on how often something happens, and a hard limit on any edge a trading bot could learn. pt-v20 passes all 40.

A rows replay 2008 and 2020, B rows count events and long-run figures over the free histories, C rows are leak and consistency checks and the cost of size, R rows cover rates, bonds and the driven 2022 market, S rows the packaged recession, D2, F1 and L1 the driven 2020 path, E1 earnings in a contraction, V1 variance over two and five years, and D1 the one-year table below. Where a row has both replays, 2008 comes before 2020.

RowWhat it checkspt-v20RealPass mark
A1Worst month's volatility in the 2008 and 2020 replays within 30% of real88.5 / 76.584.3 / 94.5Within 30%
A2Maximum drawdown in the 2008 and 2020 replays within 30% of real0.45 / 0.3680.568 / 0.339Within 30%
A3Peak stock correlation in the 2008 and 2020 replays within 0.15 of real0.784 / 0.7840.748 / 0.872Within 0.15
B1Share of sessions with the VIX above 300.0590.0821/2x to 2x
B2Mean length of a fear spell above VIX 30, sessions27221/2x to 2x
B320% bear markets per decade1.961.121/2x to 2x
B410% corrections per decade4.523.651/2x to 2x
B5Sessions down more than 5% per decade9.66.21/2x to 2x
B6Share of sessions with the VIX under 150.40.3261/2x to 2x
B7Index annual volatility, %19.118.1Within 20%
B8Long-run index return, % a year6.46.25Within 2 points of the target
B9Sd of annual index log returns, years 2-21, percent; the start-up drift: sd of the first 60 sessions' index return over the steady state's, across 50 markets16.3 / 0.79717.4 / 1Annual within 20% of 17.4; start-up 2/3x to 1.5x
C1Crash rate in years 3-21 against years 1-20.9212/3x to 1.5x
C2Histories touching the VIX ceiling, of 3000At most 1
C3Edge from reading a headline 5 ticks late, bp15.80Under 20 bp
C4a65-minute lag-1 autocorrelation of print returns, median name; Roll spread over quoted−0.016 / 1.210 / 1At or above −0.05, or Roll at most 2x quoted
C4bThe price-only rule furthest over its line on tf-suite-2026.1 (named in worst): median points over buy-and-hold, markets beaten of 200.2 / 110 / 10Every rule at most +5 points and 14 of 20
C5Stock-level (idiosyncratic) variance ratio at 60 sessions, median name, 30 histories of 2,660 sessions0.950.924In 0.80 to 1.05
C6Rank IC of the value signal on published fundamentals against the next 20 sessions: whole history, first 60 sessions0.0003 / 0.00480.0093 / 0.0093Both in −0.03 to +0.05
C7Rank IC of 12-1 and 6-1 month momentum against the next 20 sessions−0.0031 / −0.00130.0267 / 0.041312-1 in −0.04 to +0.095, 6-1 in −0.02 to +0.10
C8Lo-MacKinlay one-day loser-minus-winner book, bp a day−0.0645−1.74In −6.4 to +2.9
C9The cost of size in the agent-facing book: exponent and coefficient of the average cost against Q/V, in sigma units0.484 / 0.4240.5 / 0.5Exponent in 0.4 to 0.7 and coefficient in 0.33 to 0.67
R1Sd of the 2-year Treasury yield's daily change, bp (FRED DGS2 2015-2025)3.875.23In 3.65 to 6.80
R2Sd of the 10-year Treasury yield's daily change, bp (FRED DGS10)4.965.41In 4.54 to 6.27
R3Daily correlation of the roster index with a Treasury bond's return (minus the 10-year's change; SPY against IEF)−0.136−0.161In −0.36 to +0.03
R4Daily correlation of the roster index with an IG bond's return (minus the corporate yield's change; SPY against LQD)0.20.272In +0.15 to +0.39
E1The aggregate earnings fall around a contraction, median over the 30 histories' contractions (Shiller 1953-2020)−0.172−0.17In −0.40 to −0.046
D2The driven 2020-21 market: the index's maximum drawdown and the sessions from the pre-crash high back to it, medians over histories0.374 / 1200.339 / 126Drawdown in 0.237 to 0.441 and sessions in 63 to 252
F1The driven 2020 path's fast crash: the high to the lowest close within 60 sessions, and the sessions it took0.307 / 400.339 / 23Drop in 0.237 to 0.441 within 12 to 46 sessions
L1Look-through: sessions from the index's trough to the aggregate earnings level's trough on the driven 2020 path10.568Lead in 1 to 136 sessions
R5The driven 2022 market: the index's maximum drawdown0.2650.254Drawdown in 0.178 to 0.330
R6The driven 2022 market: the market P/E's log change per 100 bp of the corporate yield, monthly averages−4.26−5.2−10.4 to −2.6 percent
C10Timing rules on published macro data against buy-and-hold, and the drift after a published turn0.117 / 0.611 / 5 / −0.374−2 / 0.324 / −5.9 / 4.3Every rule at most +1.0 point a year and ahead in at most 2/3; drifts no larger than the S&P's after NBER turns
R7aThe equal-weight index's mean log move from the first price readable after a changed policy rate to the end of the first 65-minute bar, less the mean first bar of all days, bp: after a hike, after a cut−0.0853 / −1.730 / 0Each within 5 bp, or 2 se where wider
R7bThe audit's rate-news agent through tf.evaluate: mean annual excess over holding in points, histories ahead−0.322 / 20 / 15At most 0 points a year, ahead in at most 20 of 30
S1aThe packaged recession: share of the index's paired log fall, at its lowest, won back 252 sessions later, mean over seeds0.4940.62In 45% to 100%
S1bThe packaged recession: seeds whose cycle has left contraction and trough within 24 months of the onset3030Every seed
S2The packaged recession: the index's own rise in the 252 sessions after its low, percent, mean over seeds54.669In +25% to +80%
V1The index's variance ratio at two and five years relative to one, years 2-21 of the pooled free histories0.821 / 0.6190.93 / 0.872y/1y in 0.75 to 1.15 and 5y/1y in 0.55 to 1.20
D1Ruled bands of the one-year realism table in, on all four cellsallallEvery band in

pt-v19, the previous default, fails 16 of the 40 rows on the same pooled histories. The rows are in tf.preset_record()["long_run"].

The one-year table

The second test is a panel of statistics measured over one year, the median across 30 seeds, each against a band read from the longest real record the statistic allows. The 14 shape statistics and crisis dispersion are measured on a fixed roster of 40 companies; the four index rows on a roster that changes with each seed, because a crash rate measured on one roster describes only that roster.

StatisticWhat it checksReal bandpt-v20In band
annualised_vol_pctHow much prices move in a year12 to 4120.5yes
excess_kurtosisHow fat the tails are−13 to 2418.1yes
return_acf1Does yesterday predict today−0.07 to 0.060.013yes
abs_return_acf1Does a wild day follow a wild day0.02 to 0.170.0282yes
abs_return_acf5The same, one week apart−0.03 to 0.10.0188yes
abs_return_acf20The same, one month apart−0.05 to 0.060.0044yes
cross_sectional_corrHow much names move together0.09 to 0.490.305yes
volume_abs_return_corrDo big moves come with volume0.35 to 0.640.596yes
leverage_effectDo falls raise volatility−0.11 to 0−0.0341yes
volume_change_acf1Does volume mean-revert−0.3 to −0.2−0.268yes
corr_asymmetryDo names couple more when falling−0.15 to 0.230.0791yes
corr_asymmetry_laggedThe same, one day later−0.15 to 0.330.086yes
sector_excess_corrDo industries move together0.04 to 0.230.117yes
corr_persistence_acf1Does correlation stay high after a panic−0.48 to 0.690.23yes
crisis_sector_dispersionDoes one industry lead a crisis0.79 to 1.741.3yes
index_drift_pctWhich way the index goes1.1 to 10.37.7yes
fear_gauge_dn1How much fear a bad day buys0.39 to 3.031.66yes
fear_gauge_dn3The same, on a much worse day2.6 to 9.584.26yes
index_tail_dn3_pctHow often the index falls hard0.64 to 2.340.89yes

The table certifies runs of up to 252 trading days. At 504 days pt-v20 holds 14 of 14 graded rows of the two-year panel, but tf.envelope.check() certifies one year, and the 40 rows above are the evidence for longer runs. Some statistics pass low: volatility clustering is weaker than real at every lag, and the named gaps below say what that rules out.

Ask the package before you lean on a statistic:

import tradefloor as tf

verdict = tf.envelope.check(horizon_days=504, statistics=["abs_return_acf20"])
print(str(verdict).splitlines()[0])
OUTSIDE the envelope

The verdict then names each reason: here the horizon is past the certified 252 days, and lag-20 volatility clustering is one of the named gaps.

The named gaps

Each gap says what the model gets wrong and which uses it rules out. tf.envelope.check() refuses a question that asks for one of them.

GapWhat falls shortDo not use it for
horizonThe certified horizon is 252 daysMulti-year backtests, and anything keyed on volatility dynamics beyond one year
decay-shapeVolatility memory is weaker than real at every lagStrategies whose edge depends on volatility clustering at any lag: volatility forecasts over one to five days, and vol targeting and risk parity on a one-month or longer estimate
scenario-magnitudeA driven scenario moves prices at a quarter to a half of the real sizeSizing a scenario's impact rather than detecting it
macro-rangeThe endogenous macro state cannot reach its own crisis regimesStudying inflation regimes or policy crises from the endogenous economy alone
roster-concentrationA concentrated roster is measured on pt-v19 only, for four sector mixes and the shape rowsCiting the certification for a concentrated roster on a level or crisis row, past 504 days, on any preset but pt-v19 (the default pt-v20 included), or for a sector mix other than the four measured

These limits are measured too, and the envelope does not name them:

  • Each session opens at the last print, so there are almost no overnight gaps and a stop held overnight is safer than it would be live.
  • Nothing in the market learns your pattern and trades against you, and an order sliced over a day costs less than published studies of such orders find.
  • With the VIX held at 65 the market is 5.1 times as volatile as at a VIX of 5, against a real 6.2.
  • Some rows pass near an edge. The 2-year Treasury moves 3.87 bp a day against a real 5.23, near the floor of 3.65, and one macro timing rule uses 92% of its tolerance.

The published grade

The registration, the scripts that graded pt-v20, their inputs and the outputs of the grading run are in validation/pt-v20/ in the library repository, and validation/README.md says how to check the grade on a laptop or run it again. The one-year table is also published as data, in envelope.json, and docs/STATISTICS.md defines every statistic and gives the source of each band.

What the grade does not say

Passing means the model matches real markets on these figures and was not tuned to the exam. It does not mean a good score here predicts real returns. Some behaviors are not checked at all. Coefficients pt-v20 carries over from pt-v19 were chosen with a scoring rule over statistics that overlap the one-year table, so that table helped choose them and is weaker evidence than the rows graded on exam seeds.