Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

Case Study: CME Futures

This case study uses daily Databento data on 30 CME futures products spanning seven sectors (equity indices, treasuries, energy, metals, currencies, agriculture, and livestock) to test whether carry and term-structure signals produce tradeable alpha at a weekly cadence. Futures have a return decomposition that splits into spot and roll components, natural sector groupings that constrain diversification, and inherent leverage that magnifies both signal and friction.

The pipeline runs a long-short carry-ranked strategy with weekly Friday-close decisions and Monday-open execution, trades 30 front-month continuous contracts built with ratio back-adjustment, and prices in commission, bid-ask spread, and roll slippage. The teaching point is the distinction between average rank correlation and traded performance: portfolio Sharpe comes from magnitude at the top of the cross-section rather than from average IC, so the family that ranks best on IC and the family the selection rule carries need not be the same one. The holdout is two years of weekly decisions, a window short enough that the interval around anything it estimates is wide by construction.

At a Glance

Property Value
Asset Class CME futures (30 products, 7 sectors)
Frequency Daily data, weekly decisions
Universe 30 front-month continuous contracts
History 2011-2025
Primary Label fwd_ret_5d
CV Folds 5 (8Y train, 1Y val)
Cost Model Material (commission + spread + roll slippage)

Pipeline

Stage Notebook Chapter Description Writes
Feasibility 01_feasibility_analysis Ch6 Universe breadth per decision date, round-trip spread per product, move-to-spread scale, carry persistence, and the declared walk-forward folds none
Labels 02_labels Ch7 5-day and 21-day forward returns from ratio back-adjusted continuous data labels/fwd_ret_5d.parquet, labels/fwd_ret_21d.parquet, each with a .digest.json sidecar
Features 03_financial_features Ch8 Term structure, carry, momentum, and roll-return features features/financial.parquet
Temporal 04_model_based_features Ch9 Expanding-window ARIMA and HMM features via statsforecast features/model_based.parquet
Evaluation 05_evaluation Ch7-9 Feature-label IC diagnostics across 30 products and 7 sectors evaluation/triage_ledger.parquet, evaluation/ic_timeseries.parquet
Linear 06_linear Ch11 Ridge, LASSO, ElasticNet on carry and momentum signals Training runs and prediction sets in run_log/registry.db; coefficients under run_log/training/{hash}/, scores under run_log/predictions/{hash}/
GBM 07_gbm Ch12 LightGBM testing non-linear carry and momentum interactions Training runs and prediction sets; boosters, learning_curves.parquet, and fold_metrics.parquet under run_log/training/{hash}/
Tabular DL 08_tabular_dl Ch12 TabM rank-1 adapter MLP ensemble on flat features Training runs and prediction sets; checkpoints under run_log/training/tabular_dl/
LSTM 09_dl_lstm Ch13 Gated recurrence on the 30-product daily panel Training runs and prediction sets; checkpoints under run_log/training/deep_learning/
Latent factors (index) 10_latent_factors Ch14 Index of the two latent-factor notebooks below Nothing - it reads the registry
PCA 10a_pca Ch14 Principal components on the cross-sectional characteristics panel Training runs and prediction sets
SDF 10b_stochastic_discount_factor Ch14 Stochastic discount factor on the same panel Training runs and prediction sets
Causal DML 11_causal_dml Ch15 Does the carry signal cause future returns or proxy for risk? A row in the registry's causal_runs
Model Analysis 12_model_analysis Ch11-15 Cross-model IC comparison and fold stability diagnostics Nothing - it reads the registry
Backtest 13_backtest Ch16 Long-short carry-ranked strategy simulation One backtest run per prediction set and entry scheme; daily_returns.parquet, weights.parquet, trades.parquet, fills.parquet, equity.parquet, portfolio_state.parquet, and spec.json under run_log/backtest/{hash}/
Portfolio 14_portfolio_management Ch17 Equal-risk, score-weighted, and sector-constrained allocation One backtest run per allocation method, same artifact layout
Risk 15_risk_management Ch19 Position-level risk overlays (stop-loss, trailing stops, time exits) One backtest run per overlay variant, same artifact layout
Costs 16_costs Ch18 Commission, spread, and roll slippage impact analysis One backtest run per cost level, same artifact layout
Holdout predictions 17_holdout_predictions Ch20 Refits the resolved carrier through the holdout fold and predicts 2024-2025 One training run under a new identity whose CV declares the holdout fold, and one prediction set at split='holdout'
Holdout backtest 18_holdout_backtest Ch20 Replays the carrier's own strategy specification on the holdout prediction set One backtest run at stage='holdout', same artifact layout
Strategy Analysis 19_strategy_analysis Ch20 End-to-end strategy assessment with uncertainty-aware metrics results/strategy_assessment.json, 20_strategy_synthesis/output/cme_futures/cme_futures_tearsheet.html; rebuilds the registry's cohort_metrics and backtest_paired_metrics tables on every canonical run, pruning rows a previous selection wrote

Margin Model

Per-product margin is computed once from CME's outright maintenance rates (published on cmegroup.com) and 2025-12-31 front-month settlement prices, then expressed as a fraction of notional via ContractSpec.margin_pct = (initial, maintenance) in data/futures/market/futures_specs.yaml. The engine applies the ratio to each historical bar's notional, so the dollar margin moves with price even though the rate is anchored at one point. For the 8 products not covered by the CSV (the 4 equity-index e-minis ES/NQ/YM/RTY and 4 energies CL/NG/HO/RB), per-category SPAN-style initial-margin approximations are used (initial: 5% equity_index, 8% energy; the table reports maintenance, derived as initial ÷ 1.10 per the CME SPAN convention).

This is a stable-pct approximation. CME publishes maintenance dollars in scan-volume steps that adjust roughly with realized volatility. Historical SPAN snapshots (free up to ~5 years from CME's historical-margins page and paid beyond that via CME Datamine catalog F001) would let us anchor period-specific ratios; the marginal effect on conclusions for 30 liquid CME products is small but non-zero. The table below shows representative pct drift between start- and end-of-window prices for products spanning the secular regimes:

Product StartDate StartPx EndDate EndPx Ratio Anchored pct Start-window implied pct Drift
ES 2011-01-03 1,103 2025-12-31 6,942 6.30× 4.54% 28.61% +530%
NQ 2011-01-03 2,612 2025-12-31 25,657 9.82× 4.54% 44.64% +882%
GC 2011-01-03 1,978 2025-12-31 4,351 2.20× 6.28% 13.82% +120%
ZN 2011-01-03 100.9 2025-12-31 112.6 1.12× 1.67% 1.86% +12%
CL 2011-01-03 189.6 2025-12-31 57.9 0.31× 7.27% 2.22% −69%
ZC 2011-01-03 723.2 2025-12-30 440.5 0.61× 4.43% 2.70% −39%
NG 2011-01-03 221.3 2025-12-31 3.97 0.018× 7.27% 0.13% −98%

Anchored pct is closest to truth near 2025-12-31; back in 2011 the stable-pct approximation under-margins high-momentum equity indices (engine accepts orders a live broker may have rejected) and over-margins crashed commodities (engine rejects orders a live broker may have accepted). The effect is confined to absolute levels in the 2011-2015 portion of the validation window; relative comparisons across families, allocators, and cost variants are unaffected. Holdout (2024–2025) is anchored at the same window as the pct calculation, so this drift does not bear on the holdout.

Account sizing: initial_cash is $10M in config/setup.yaml. The 30-product universe spans contract notionals from ≈$35k (NG) to ≈$22M (ZT 2-year T-Note), and the engine sizes positions as target_notional / contract_notional → integer contracts. At k=5 per side (10 positions total) the per-position dollar budget is 5% × cash; $10M clears the binding constraint (ES at ≈$347k holdout notional in the late window) and lets all 30 products participate. NQ in the very-late holdout (Dec 2025, peak notional ≈$513k vs $500k per-position budget) is the one residual integer-share footnote.

Running

# From repo root
uv run python case_studies/cme_futures/01_feasibility_analysis.py
uv run python case_studies/cme_futures/02_labels.py
uv run python case_studies/cme_futures/03_financial_features.py
uv run python case_studies/cme_futures/04_model_based_features.py
uv run python case_studies/cme_futures/05_evaluation.py
uv run python case_studies/cme_futures/06_linear.py
uv run python case_studies/cme_futures/07_gbm.py
uv run python case_studies/cme_futures/08_tabular_dl.py
uv run python case_studies/cme_futures/09_dl_lstm.py
uv run python case_studies/cme_futures/10a_pca.py
uv run python case_studies/cme_futures/10b_stochastic_discount_factor.py
uv run python case_studies/cme_futures/10_latent_factors.py   # summarizes 10a-10b
uv run python case_studies/cme_futures/11_causal_dml.py
uv run python case_studies/cme_futures/12_model_analysis.py
uv run python case_studies/cme_futures/13_backtest.py
uv run python case_studies/cme_futures/14_portfolio_management.py
uv run python case_studies/cme_futures/15_risk_management.py
uv run python case_studies/cme_futures/16_costs.py
uv run python case_studies/cme_futures/17_holdout_predictions.py
uv run python case_studies/cme_futures/18_holdout_backtest.py
uv run python case_studies/cme_futures/19_strategy_analysis.py

Results

This README describes how the case study is built, not what it found. Results are not restated here: the registry is rebuilt whenever the case study is re-derived, and a number copied into prose stays correct only until the next rebuild.

19_strategy_analysis reads the registry back and reports the selected configuration with its interval evidence. That notebook, and the registry it reads, are where a result comes from.

To read the results without training anything, download the published bundle, which carries the registry and the artifacts behind it:

uv run python scripts/download_artifacts.py --cs cme_futures

Two bundles are published, and they are separate generations rather than revisions of one another. v3.1.0-artifacts is current and is what the command above fetches. v3.0.0-artifacts holds the results as first published. The 3.1 rebuild re-keyed every content-addressed hash, so a hash taken from one bundle does not resolve in the other.

Run Log

run_log/registry.db records every training run, prediction set and backtest, each addressed by a hash of the specification that produced it. The artifacts sit beside it under run_log/training/, run_log/predictions/ and run_log/backtest/.