How it is measured
Every check and its pass mark is written down and committed before the grade, the grade runs on seeds the tuning never used, and the scripts and their output are published. pt-v20 passed all 40 of its registered rows that way, and all 19 statistics of the one-year table are inside the ranges real markets show.
Registration, then one grade
The seeds
For pt-v20 the final settings were chosen on practice seeds 201 to 230, 501 to 530 and 801 to 830. The grade ran on exam seeds 101 to 130, 401 to 430 and 701 to 730: 90 histories of 21 years with nothing imposed, plus replays of 2008 and 2020 with the real VIX, the real 2020-21 and 2022 economies driven through the model, the packaged recession, and the leak-finding strategies.
The 40 registered rows
Each row is something a user would notice, with a tolerance that is easy to read: within 30% on the size of a crash, half to twice the real rate on how often something happens, and a hard limit on any edge a trading bot could learn. pt-v20 passes all 40.
A rows replay 2008 and 2020, B rows count events and long-run figures over the free histories, C rows are leak and consistency checks and the cost of size, R rows cover rates, bonds and the driven 2022 market, S rows the packaged recession, D2, F1 and L1 the driven 2020 path, E1 earnings in a contraction, V1 variance over two and five years, and D1 the one-year table below. Where a row has both replays, 2008 comes before 2020.
| Row | What it checks | pt-v20 | Real | Pass mark |
|---|---|---|---|---|
A1 | Worst month's volatility in the 2008 and 2020 replays within 30% of real | 88.5 / 76.5 | 84.3 / 94.5 | Within 30% |
A2 | Maximum drawdown in the 2008 and 2020 replays within 30% of real | 0.45 / 0.368 | 0.568 / 0.339 | Within 30% |
A3 | Peak stock correlation in the 2008 and 2020 replays within 0.15 of real | 0.784 / 0.784 | 0.748 / 0.872 | Within 0.15 |
B1 | Share of sessions with the VIX above 30 | 0.059 | 0.082 | 1/2x to 2x |
B2 | Mean length of a fear spell above VIX 30, sessions | 27 | 22 | 1/2x to 2x |
B3 | 20% bear markets per decade | 1.96 | 1.12 | 1/2x to 2x |
B4 | 10% corrections per decade | 4.52 | 3.65 | 1/2x to 2x |
B5 | Sessions down more than 5% per decade | 9.6 | 6.2 | 1/2x to 2x |
B6 | Share of sessions with the VIX under 15 | 0.4 | 0.326 | 1/2x to 2x |
B7 | Index annual volatility, % | 19.1 | 18.1 | Within 20% |
B8 | Long-run index return, % a year | 6.4 | 6.25 | Within 2 points of the target |
B9 | Sd of annual index log returns, years 2-21, percent; the start-up drift: sd of the first 60 sessions' index return over the steady state's, across 50 markets | 16.3 / 0.797 | 17.4 / 1 | Annual within 20% of 17.4; start-up 2/3x to 1.5x |
C1 | Crash rate in years 3-21 against years 1-2 | 0.92 | 1 | 2/3x to 1.5x |
C2 | Histories touching the VIX ceiling, of 30 | 0 | 0 | At most 1 |
C3 | Edge from reading a headline 5 ticks late, bp | 15.8 | 0 | Under 20 bp |
C4a | 65-minute lag-1 autocorrelation of print returns, median name; Roll spread over quoted | −0.016 / 1.21 | 0 / 1 | At or above −0.05, or Roll at most 2x quoted |
C4b | The price-only rule furthest over its line on tf-suite-2026.1 (named in worst): median points over buy-and-hold, markets beaten of 20 | 0.2 / 11 | 0 / 10 | Every rule at most +5 points and 14 of 20 |
C5 | Stock-level (idiosyncratic) variance ratio at 60 sessions, median name, 30 histories of 2,660 sessions | 0.95 | 0.924 | In 0.80 to 1.05 |
C6 | Rank IC of the value signal on published fundamentals against the next 20 sessions: whole history, first 60 sessions | 0.0003 / 0.0048 | 0.0093 / 0.0093 | Both in −0.03 to +0.05 |
C7 | Rank IC of 12-1 and 6-1 month momentum against the next 20 sessions | −0.0031 / −0.0013 | 0.0267 / 0.0413 | 12-1 in −0.04 to +0.095, 6-1 in −0.02 to +0.10 |
C8 | Lo-MacKinlay one-day loser-minus-winner book, bp a day | −0.0645 | −1.74 | In −6.4 to +2.9 |
C9 | The cost of size in the agent-facing book: exponent and coefficient of the average cost against Q/V, in sigma units | 0.484 / 0.424 | 0.5 / 0.5 | Exponent in 0.4 to 0.7 and coefficient in 0.33 to 0.67 |
R1 | Sd of the 2-year Treasury yield's daily change, bp (FRED DGS2 2015-2025) | 3.87 | 5.23 | In 3.65 to 6.80 |
R2 | Sd of the 10-year Treasury yield's daily change, bp (FRED DGS10) | 4.96 | 5.41 | In 4.54 to 6.27 |
R3 | Daily correlation of the roster index with a Treasury bond's return (minus the 10-year's change; SPY against IEF) | −0.136 | −0.161 | In −0.36 to +0.03 |
R4 | Daily correlation of the roster index with an IG bond's return (minus the corporate yield's change; SPY against LQD) | 0.2 | 0.272 | In +0.15 to +0.39 |
E1 | The aggregate earnings fall around a contraction, median over the 30 histories' contractions (Shiller 1953-2020) | −0.172 | −0.17 | In −0.40 to −0.046 |
D2 | The driven 2020-21 market: the index's maximum drawdown and the sessions from the pre-crash high back to it, medians over histories | 0.374 / 120 | 0.339 / 126 | Drawdown in 0.237 to 0.441 and sessions in 63 to 252 |
F1 | The driven 2020 path's fast crash: the high to the lowest close within 60 sessions, and the sessions it took | 0.307 / 40 | 0.339 / 23 | Drop in 0.237 to 0.441 within 12 to 46 sessions |
L1 | Look-through: sessions from the index's trough to the aggregate earnings level's trough on the driven 2020 path | 10.5 | 68 | Lead in 1 to 136 sessions |
R5 | The driven 2022 market: the index's maximum drawdown | 0.265 | 0.254 | Drawdown in 0.178 to 0.330 |
R6 | The driven 2022 market: the market P/E's log change per 100 bp of the corporate yield, monthly averages | −4.26 | −5.2 | −10.4 to −2.6 percent |
C10 | Timing rules on published macro data against buy-and-hold, and the drift after a published turn | 0.117 / 0.611 / 5 / −0.374 | −2 / 0.324 / −5.9 / 4.3 | Every rule at most +1.0 point a year and ahead in at most 2/3; drifts no larger than the S&P's after NBER turns |
R7a | The equal-weight index's mean log move from the first price readable after a changed policy rate to the end of the first 65-minute bar, less the mean first bar of all days, bp: after a hike, after a cut | −0.0853 / −1.73 | 0 / 0 | Each within 5 bp, or 2 se where wider |
R7b | The audit's rate-news agent through tf.evaluate: mean annual excess over holding in points, histories ahead | −0.322 / 2 | 0 / 15 | At most 0 points a year, ahead in at most 20 of 30 |
S1a | The packaged recession: share of the index's paired log fall, at its lowest, won back 252 sessions later, mean over seeds | 0.494 | 0.62 | In 45% to 100% |
S1b | The packaged recession: seeds whose cycle has left contraction and trough within 24 months of the onset | 30 | 30 | Every seed |
S2 | The packaged recession: the index's own rise in the 252 sessions after its low, percent, mean over seeds | 54.6 | 69 | In +25% to +80% |
V1 | The index's variance ratio at two and five years relative to one, years 2-21 of the pooled free histories | 0.821 / 0.619 | 0.93 / 0.87 | 2y/1y in 0.75 to 1.15 and 5y/1y in 0.55 to 1.20 |
D1 | Ruled bands of the one-year realism table in, on all four cells | all | all | Every band in |
pt-v19, the previous default, fails 16 of the 40 rows on the same pooled histories. The rows are in tf.preset_record()["long_run"].
The one-year table
The second test is a panel of statistics measured over one year, the median across 30 seeds, each against a band read from the longest real record the statistic allows. The 14 shape statistics and crisis dispersion are measured on a fixed roster of 40 companies; the four index rows on a roster that changes with each seed, because a crash rate measured on one roster describes only that roster.
| Statistic | What it checks | Real band | pt-v20 | In band |
|---|---|---|---|---|
annualised_vol_pct | How much prices move in a year | 12 to 41 | 20.5 | yes |
excess_kurtosis | How fat the tails are | −13 to 24 | 18.1 | yes |
return_acf1 | Does yesterday predict today | −0.07 to 0.06 | 0.013 | yes |
abs_return_acf1 | Does a wild day follow a wild day | 0.02 to 0.17 | 0.0282 | yes |
abs_return_acf5 | The same, one week apart | −0.03 to 0.1 | 0.0188 | yes |
abs_return_acf20 | The same, one month apart | −0.05 to 0.06 | 0.0044 | yes |
cross_sectional_corr | How much names move together | 0.09 to 0.49 | 0.305 | yes |
volume_abs_return_corr | Do big moves come with volume | 0.35 to 0.64 | 0.596 | yes |
leverage_effect | Do falls raise volatility | −0.11 to 0 | −0.0341 | yes |
volume_change_acf1 | Does volume mean-revert | −0.3 to −0.2 | −0.268 | yes |
corr_asymmetry | Do names couple more when falling | −0.15 to 0.23 | 0.0791 | yes |
corr_asymmetry_lagged | The same, one day later | −0.15 to 0.33 | 0.086 | yes |
sector_excess_corr | Do industries move together | 0.04 to 0.23 | 0.117 | yes |
corr_persistence_acf1 | Does correlation stay high after a panic | −0.48 to 0.69 | 0.23 | yes |
crisis_sector_dispersion | Does one industry lead a crisis | 0.79 to 1.74 | 1.3 | yes |
index_drift_pct | Which way the index goes | 1.1 to 10.3 | 7.7 | yes |
fear_gauge_dn1 | How much fear a bad day buys | 0.39 to 3.03 | 1.66 | yes |
fear_gauge_dn3 | The same, on a much worse day | 2.6 to 9.58 | 4.26 | yes |
index_tail_dn3_pct | How often the index falls hard | 0.64 to 2.34 | 0.89 | yes |
The table certifies runs of up to 252 trading days. At 504 days pt-v20 holds 14 of 14 graded rows of the two-year panel, but tf.envelope.check() certifies one year, and the 40 rows above are the evidence for longer runs. Some statistics pass low: volatility clustering is weaker than real at every lag, and the named gaps below say what that rules out.
Ask the package before you lean on a statistic:
import tradefloor as tf verdict = tf.envelope.check(horizon_days=504, statistics=["abs_return_acf20"]) print(str(verdict).splitlines()[0])
OUTSIDE the envelope
The verdict then names each reason: here the horizon is past the certified 252 days, and lag-20 volatility clustering is one of the named gaps.
The named gaps
Each gap says what the model gets wrong and which uses it rules out. tf.envelope.check() refuses a question that asks for one of them.
| Gap | What falls short | Do not use it for |
|---|---|---|
horizon | The certified horizon is 252 days | Multi-year backtests, and anything keyed on volatility dynamics beyond one year |
decay-shape | Volatility memory is weaker than real at every lag | Strategies whose edge depends on volatility clustering at any lag: volatility forecasts over one to five days, and vol targeting and risk parity on a one-month or longer estimate |
scenario-magnitude | A driven scenario moves prices at a quarter to a half of the real size | Sizing a scenario's impact rather than detecting it |
macro-range | The endogenous macro state cannot reach its own crisis regimes | Studying inflation regimes or policy crises from the endogenous economy alone |
roster-concentration | A concentrated roster is measured on pt-v19 only, for four sector mixes and the shape rows | Citing the certification for a concentrated roster on a level or crisis row, past 504 days, on any preset but pt-v19 (the default pt-v20 included), or for a sector mix other than the four measured |
These limits are measured too, and the envelope does not name them:
- Each session opens at the last print, so there are almost no overnight gaps and a stop held overnight is safer than it would be live.
- Nothing in the market learns your pattern and trades against you, and an order sliced over a day costs less than published studies of such orders find.
- With the VIX held at 65 the market is 5.1 times as volatile as at a VIX of 5, against a real 6.2.
- Some rows pass near an edge. The 2-year Treasury moves 3.87 bp a day against a real 5.23, near the floor of 3.65, and one macro timing rule uses 92% of its tolerance.
The published grade
The registration, the scripts that graded pt-v20, their inputs and the outputs of the grading run are in validation/pt-v20/ in the library repository, and validation/README.md says how to check the grade on a laptop or run it again. The one-year table is also published as data, in envelope.json, and docs/STATISTICS.md defines every statistic and gives the source of each band.
What the grade does not say
Passing means the model matches real markets on these figures and was not tuned to the exam. It does not mean a good score here predicts real returns. Some behaviors are not checked at all. Coefficients pt-v20 carries over from pt-v19 were chosen with a scoring rule over statistics that overlap the one-year table, so that table helped choose them and is weaker evidence than the rows graded on exam seeds.