26 claimsValidated numbers
Every specific figure this documentation prints was measured rather than guessed. This is where each one is held to account.
A claim with no check behind it drifts from the code eventually, and the drift is silent: the number goes on reading as authoritative while the software has moved. So each row names the test that measures its figure, and a test in this repository asserts that every one of those tests exists and that the figure appears in its source. A number that changes in the code and not on the page fails a build.
What this does not catch is a new claim added to the documentation and never added to the table. That is why the table renders here rather than living only in the test suite: a figure that is not in this list is visibly not in this list.
| Figure | What it measures | Checked by |
|---|---|---|
| 869 | The total implementation shortfall of the worked order, which is the same under either delay basis. | tests/test_execution.py test_the_delay_basis_reattributes_cost_without_changing_the_total |
| 200.0 | Delay cost of that order charged on the quantity ordered. | tests/test_execution.py test_the_delay_basis_reattributes_cost_without_changing_the_total |
| 180.0 | Delay cost of the same order charged on the quantity executed — the 20 units that move into opportunity cost. | tests/test_execution.py test_the_delay_basis_reattributes_cost_without_changing_the_total |
| 0.0246 | Fraction by which permanent impact shortens the execution half-life at a period length of 1. | tests/test_execution.py test_permanent_impact_leaves_the_schedule_alone_only_in_the_limit |
| 0.00025 | The same fraction at a period length of 0.01, which is the O(tau) behaviour the algebra predicts. | tests/test_execution.py test_permanent_impact_leaves_the_schedule_alone_only_in_the_limit |
| 0.0625 | Fixed cost per share that raises the expected cost by exactly itself times the quantity, leaving the schedule unmoved. | tests/test_execution.py test_a_fixed_cost_moves_the_cost_by_itself_and_leaves_the_schedule_alone |
| 200_000.0 | Each slice of the risk-neutral schedule for a million shares over five periods: TWAP, exactly. | tests/test_execution.py test_risk_neutrality_is_exactly_a_straight_line |
| 3.0 | Effective number of trials among forty driven by one common factor, by the eigenvalue method. | tests/test_validation.py test_one_factor_trials_collapse_to_a_handful |
| 40.0 | Effective number of trials among forty independent ones: the raw count, which is what makes the correction meaningful. | tests/test_validation.py test_independent_trials_are_worth_their_raw_count |
| 0.8244 | Deflated Sharpe ratio of those correlated trials on the effective count. | tests/test_validation.py test_the_effective_count_matters_exactly_where_trials_are_correlated |
| 0.7196 | The same figure deflated against the raw count of forty — ten points of probability the raw count throws away. | tests/test_validation.py test_the_effective_count_matters_exactly_where_trials_are_correlated |
| 2.71 | Years of daily data before an annualised Sharpe ratio of 1.0 is distinguishable from zero at 95% confidence. | tests/test_validation.py test_a_bigger_sharpe_needs_a_shorter_record |
| 87.124248 | Total cost of the worked order in basis points of its paper notional. | tests/test_execution_over_the_wire.py test_the_decomposition_survives_the_wire |
| 1.64724 | Half-life of the worked execution schedule at a risk aversion of 1e-6, in periods. | tests/test_execution_over_the_wire.py test_the_schedule_survives_the_wire |
| 140_000 | Budget on the whole tool listing, in characters of serialised JSON. | tests/test_skill.py test_the_tool_listing_fits_its_budget |
| 12_000 | Budget on any single tool, in characters of serialised JSON. | tests/test_skill.py test_no_single_tool_takes_more_than_its_share |
| 279.0 | Holding-period return of a 3% 2031 bond over one year on the bundled quote screen, in basis points. Quoted in the bond_carry_rolldown description, which is asserted in the same test. | tests/test_model_validation.py test_the_figure_the_description_quotes_is_the_one_it_computes |
| 2.49 | Roll-down inside that 279 basis points, per 100 of face. The financing cost is the other 0.50, so nearly all of the return is the curve failing to evolve to its forwards. | tests/test_model_validation.py test_the_figure_the_description_quotes_is_the_one_it_computes |
| -0.245 | The Acerbi-Szekely conditional statistic when the true volatility is double the forecast. Quoted in validate_risk_model's note to warn that the statistic is not on the scale of the error it detects. | tests/test_model_validation.py test_the_note_gives_the_measured_scale_of_the_shortfall_statistic |
| 10.00% of the time | How often a breach follows a breach under a regime-switching volatility, against 1.84% after a calm day — the clustering a breach count cannot see. | tests/test_model_validation.py test_clustered_breaches_are_caught_by_independence_not_by_the_count |
| 28.2 to 22.45 | The 99% breach count over 2000 observations of a regime-switching series, under normal innovations and under an estimated tail, against a nominal 20. Quoted in conditional_volatility's note to say what estimating the tail buys and what it leaves behind. | tests/test_model_validation.py test_the_note_refuses_the_flattering_summary |
| 2.5, 11.0, 8.2 and 0.6 | The gap between the simulated horizon value at risk and the substitution it replaces — the horizon volatility times the one-step quantile multiplier — in units of the simulation's own standard error, over four samples. Quoted to say that the gap is real on average and not decisive on every series. | tests/test_model_validation.py test_the_horizon_quantile_is_not_the_volatility_times_a_multiplier |
| 41% larger at four degrees of freedom | How much wider a forecast becomes if the raw Student-t quantile is used where the standardised one belongs. The reason conditional_volatility returns the multiplier rather than describing how to build it. | tests/test_model_validation.py test_the_quantile_multiplier_is_the_standardised_quantile |
| 0.25 | The tail index of a Student-t on four degrees of freedom, which is the reciprocal of its degrees of freedom. The known truth the extreme-value method's fitted shape is measured against, rather than against its own output. | tests/test_risk.py test_the_fitted_tail_reports_the_fit_and_not_only_the_figure |
| 99.99% | The confidence at which historical simulation on 2,000 observations can only return the worst loss observed, and the fitted tail exceeds it. At 99% the two agree to within a fifth, which is the other half of the claim: the fitted method is not simply wider everywhere. | tests/test_risk.py test_the_fitted_tail_exceeds_the_historical_one_far_out |
| 6.22 | Copula degrees of freedom fitted over the wire to 700 observations of a shared Student-t factor built at four, biased high because the idiosyncratic noise dilutes the mixing variable the dependence rides on. | tests/test_risk_over_the_wire.py test_the_copula_over_the_wire |
Claims that were wrong when measured
The table above exists because several of these were. Each of the following was written down as plausible, measured afterwards, and corrected.
- Permanent impact was said not to enter an optimal execution schedule. In
continuous time it does not; in discrete time it acts through
eta - gamma·tau/2, shortening the half-life by 2.5% at a period length of 1. - The bias of uncorrected excess kurtosis on normal data is
-6/(n+1), not the commonly quoted-6/n— four standard errors apart at twenty observations. - A futures convexity adjustment was quoted at 5 basis points and measured at 51.
- “More trials never make the evidence stronger” is true of the deflation and false of the tool that applies it: adding columns changes which trial is selected and what the trials' variance is, so the deflated figure legitimately moves in either direction.