26 claimsValidated numbers

Every specific figure this documentation prints was measured rather than guessed. This is where each one is held to account.

A claim with no check behind it drifts from the code eventually, and the drift is silent: the number goes on reading as authoritative while the software has moved. So each row names the test that measures its figure, and a test in this repository asserts that every one of those tests exists and that the figure appears in its source. A number that changes in the code and not on the page fails a build.

What this does not catch is a new claim added to the documentation and never added to the table. That is why the table renders here rather than living only in the test suite: a figure that is not in this list is visibly not in this list.

FigureWhat it measuresChecked by
869The total implementation shortfall of the worked order, which is the same under either delay basis.tests/test_execution.py
test_the_delay_basis_reattributes_cost_without_changing_the_total
200.0Delay cost of that order charged on the quantity ordered.tests/test_execution.py
test_the_delay_basis_reattributes_cost_without_changing_the_total
180.0Delay cost of the same order charged on the quantity executed — the 20 units that move into opportunity cost.tests/test_execution.py
test_the_delay_basis_reattributes_cost_without_changing_the_total
0.0246Fraction by which permanent impact shortens the execution half-life at a period length of 1.tests/test_execution.py
test_permanent_impact_leaves_the_schedule_alone_only_in_the_limit
0.00025The same fraction at a period length of 0.01, which is the O(tau) behaviour the algebra predicts.tests/test_execution.py
test_permanent_impact_leaves_the_schedule_alone_only_in_the_limit
0.0625Fixed cost per share that raises the expected cost by exactly itself times the quantity, leaving the schedule unmoved.tests/test_execution.py
test_a_fixed_cost_moves_the_cost_by_itself_and_leaves_the_schedule_alone
200_000.0Each slice of the risk-neutral schedule for a million shares over five periods: TWAP, exactly.tests/test_execution.py
test_risk_neutrality_is_exactly_a_straight_line
3.0Effective number of trials among forty driven by one common factor, by the eigenvalue method.tests/test_validation.py
test_one_factor_trials_collapse_to_a_handful
40.0Effective number of trials among forty independent ones: the raw count, which is what makes the correction meaningful.tests/test_validation.py
test_independent_trials_are_worth_their_raw_count
0.8244Deflated Sharpe ratio of those correlated trials on the effective count.tests/test_validation.py
test_the_effective_count_matters_exactly_where_trials_are_correlated
0.7196The same figure deflated against the raw count of forty — ten points of probability the raw count throws away.tests/test_validation.py
test_the_effective_count_matters_exactly_where_trials_are_correlated
2.71Years of daily data before an annualised Sharpe ratio of 1.0 is distinguishable from zero at 95% confidence.tests/test_validation.py
test_a_bigger_sharpe_needs_a_shorter_record
87.124248Total cost of the worked order in basis points of its paper notional.tests/test_execution_over_the_wire.py
test_the_decomposition_survives_the_wire
1.64724Half-life of the worked execution schedule at a risk aversion of 1e-6, in periods.tests/test_execution_over_the_wire.py
test_the_schedule_survives_the_wire
140_000Budget on the whole tool listing, in characters of serialised JSON.tests/test_skill.py
test_the_tool_listing_fits_its_budget
12_000Budget on any single tool, in characters of serialised JSON.tests/test_skill.py
test_no_single_tool_takes_more_than_its_share
279.0Holding-period return of a 3% 2031 bond over one year on the bundled quote screen, in basis points. Quoted in the bond_carry_rolldown description, which is asserted in the same test.tests/test_model_validation.py
test_the_figure_the_description_quotes_is_the_one_it_computes
2.49Roll-down inside that 279 basis points, per 100 of face. The financing cost is the other 0.50, so nearly all of the return is the curve failing to evolve to its forwards.tests/test_model_validation.py
test_the_figure_the_description_quotes_is_the_one_it_computes
-0.245The Acerbi-Szekely conditional statistic when the true volatility is double the forecast. Quoted in validate_risk_model's note to warn that the statistic is not on the scale of the error it detects.tests/test_model_validation.py
test_the_note_gives_the_measured_scale_of_the_shortfall_statistic
10.00% of the timeHow often a breach follows a breach under a regime-switching volatility, against 1.84% after a calm day — the clustering a breach count cannot see.tests/test_model_validation.py
test_clustered_breaches_are_caught_by_independence_not_by_the_count
28.2 to 22.45The 99% breach count over 2000 observations of a regime-switching series, under normal innovations and under an estimated tail, against a nominal 20. Quoted in conditional_volatility's note to say what estimating the tail buys and what it leaves behind.tests/test_model_validation.py
test_the_note_refuses_the_flattering_summary
2.5, 11.0, 8.2 and 0.6The gap between the simulated horizon value at risk and the substitution it replaces — the horizon volatility times the one-step quantile multiplier — in units of the simulation's own standard error, over four samples. Quoted to say that the gap is real on average and not decisive on every series.tests/test_model_validation.py
test_the_horizon_quantile_is_not_the_volatility_times_a_multiplier
41% larger at four degrees of freedomHow much wider a forecast becomes if the raw Student-t quantile is used where the standardised one belongs. The reason conditional_volatility returns the multiplier rather than describing how to build it.tests/test_model_validation.py
test_the_quantile_multiplier_is_the_standardised_quantile
0.25The tail index of a Student-t on four degrees of freedom, which is the reciprocal of its degrees of freedom. The known truth the extreme-value method's fitted shape is measured against, rather than against its own output.tests/test_risk.py
test_the_fitted_tail_reports_the_fit_and_not_only_the_figure
99.99%The confidence at which historical simulation on 2,000 observations can only return the worst loss observed, and the fitted tail exceeds it. At 99% the two agree to within a fifth, which is the other half of the claim: the fitted method is not simply wider everywhere.tests/test_risk.py
test_the_fitted_tail_exceeds_the_historical_one_far_out
6.22Copula degrees of freedom fitted over the wire to 700 observations of a shared Student-t factor built at four, biased high because the idiosyncratic noise dilutes the mixing variable the dependence rides on.tests/test_risk_over_the_wire.py
test_the_copula_over_the_wire

Claims that were wrong when measured

The table above exists because several of these were. Each of the following was written down as plausible, measured afterwards, and corrected.