This case study uses daily Databento data on 30 CME futures products spanning seven sectors (equity indices, treasuries, energy, metals, currencies, agriculture, and livestock) to test whether carry and term-structure signals produce tradeable alpha at a weekly cadence. Futures have a return decomposition that splits into spot and roll components, natural sector groupings that constrain diversification, and inherent leverage that magnifies both signal and friction.
The pipeline runs a long-short carry-ranked strategy with weekly Friday-close decisions and Monday-open execution, trades 30 front-month continuous contracts built with ratio back-adjustment, and prices in commission, bid-ask spread, and roll slippage. The teaching point is the distinction between average rank correlation and traded performance: portfolio Sharpe comes from magnitude at the top of the cross-section rather than from average IC, so the family that ranks best on IC and the family the selection rule carries need not be the same one. The holdout is two years of weekly decisions, a window short enough that the interval around anything it estimates is wide by construction.
| Property | Value |
|---|---|
| Asset Class | CME futures (30 products, 7 sectors) |
| Frequency | Daily data, weekly decisions |
| Universe | 30 front-month continuous contracts |
| History | 2011-2025 |
| Primary Label | fwd_ret_5d |
| CV Folds | 5 (8Y train, 1Y val) |
| Cost Model | Material (commission + spread + roll slippage) |
| Stage | Notebook | Chapter | Description | Writes |
|---|---|---|---|---|
| Feasibility | 01_feasibility_analysis |
Ch6 | Universe breadth per decision date, round-trip spread per product, move-to-spread scale, carry persistence, and the declared walk-forward folds | none |
| Labels | 02_labels |
Ch7 | 5-day and 21-day forward returns from ratio back-adjusted continuous data | labels/fwd_ret_5d.parquet, labels/fwd_ret_21d.parquet, each with a .digest.json sidecar |
| Features | 03_financial_features |
Ch8 | Term structure, carry, momentum, and roll-return features | features/financial.parquet |
| Temporal | 04_model_based_features |
Ch9 | Expanding-window ARIMA and HMM features via statsforecast | features/model_based.parquet |
| Evaluation | 05_evaluation |
Ch7-9 | Feature-label IC diagnostics across 30 products and 7 sectors | evaluation/triage_ledger.parquet, evaluation/ic_timeseries.parquet |
| Linear | 06_linear |
Ch11 | Ridge, LASSO, ElasticNet on carry and momentum signals | Training runs and prediction sets in run_log/registry.db; coefficients under run_log/training/{hash}/, scores under run_log/predictions/{hash}/ |
| GBM | 07_gbm |
Ch12 | LightGBM testing non-linear carry and momentum interactions | Training runs and prediction sets; boosters, learning_curves.parquet, and fold_metrics.parquet under run_log/training/{hash}/ |
| Tabular DL | 08_tabular_dl |
Ch12 | TabM rank-1 adapter MLP ensemble on flat features | Training runs and prediction sets; checkpoints under run_log/training/tabular_dl/ |
| LSTM | 09_dl_lstm |
Ch13 | Gated recurrence on the 30-product daily panel | Training runs and prediction sets; checkpoints under run_log/training/deep_learning/ |
| Latent factors (index) | 10_latent_factors |
Ch14 | Index of the two latent-factor notebooks below | Nothing - it reads the registry |
| PCA | 10a_pca |
Ch14 | Principal components on the cross-sectional characteristics panel | Training runs and prediction sets |
| SDF | 10b_stochastic_discount_factor |
Ch14 | Stochastic discount factor on the same panel | Training runs and prediction sets |
| Causal DML | 11_causal_dml |
Ch15 | Does the carry signal cause future returns or proxy for risk? | A row in the registry's causal_runs |
| Model Analysis | 12_model_analysis |
Ch11-15 | Cross-model IC comparison and fold stability diagnostics | Nothing - it reads the registry |
| Backtest | 13_backtest |
Ch16 | Long-short carry-ranked strategy simulation | One backtest run per prediction set and entry scheme; daily_returns.parquet, weights.parquet, trades.parquet, fills.parquet, equity.parquet, portfolio_state.parquet, and spec.json under run_log/backtest/{hash}/ |
| Portfolio | 14_portfolio_management |
Ch17 | Equal-risk, score-weighted, and sector-constrained allocation | One backtest run per allocation method, same artifact layout |
| Risk | 15_risk_management |
Ch19 | Position-level risk overlays (stop-loss, trailing stops, time exits) | One backtest run per overlay variant, same artifact layout |
| Costs | 16_costs |
Ch18 | Commission, spread, and roll slippage impact analysis | One backtest run per cost level, same artifact layout |
| Holdout predictions | 17_holdout_predictions |
Ch20 | Refits the resolved carrier through the holdout fold and predicts 2024-2025 | One training run under a new identity whose CV declares the holdout fold, and one prediction set at split='holdout' |
| Holdout backtest | 18_holdout_backtest |
Ch20 | Replays the carrier's own strategy specification on the holdout prediction set | One backtest run at stage='holdout', same artifact layout |
| Strategy Analysis | 19_strategy_analysis |
Ch20 | End-to-end strategy assessment with uncertainty-aware metrics | results/strategy_assessment.json, 20_strategy_synthesis/output/cme_futures/cme_futures_tearsheet.html; rebuilds the registry's cohort_metrics and backtest_paired_metrics tables on every canonical run, pruning rows a previous selection wrote |
Per-product margin is computed once from CME's outright maintenance rates (published on cmegroup.com) and 2025-12-31 front-month settlement prices, then expressed as a fraction of notional via ContractSpec.margin_pct = (initial, maintenance) in data/futures/market/futures_specs.yaml. The engine applies the ratio to each historical bar's notional, so the dollar margin moves with price even though the rate is anchored at one point. For the 8 products not covered by the CSV (the 4 equity-index e-minis ES/NQ/YM/RTY and 4 energies CL/NG/HO/RB), per-category SPAN-style initial-margin approximations are used (initial: 5% equity_index, 8% energy; the table reports maintenance, derived as initial ÷ 1.10 per the CME SPAN convention).
This is a stable-pct approximation. CME publishes maintenance dollars in scan-volume steps that adjust roughly with realized volatility. Historical SPAN snapshots (free up to ~5 years from CME's historical-margins page and paid beyond that via CME Datamine catalog F001) would let us anchor period-specific ratios; the marginal effect on conclusions for 30 liquid CME products is small but non-zero. The table below shows representative pct drift between start- and end-of-window prices for products spanning the secular regimes:
| Product | StartDate | StartPx | EndDate | EndPx | Ratio | Anchored pct | Start-window implied pct | Drift |
|---|---|---|---|---|---|---|---|---|
| ES | 2011-01-03 | 1,103 | 2025-12-31 | 6,942 | 6.30× | 4.54% | 28.61% | +530% |
| NQ | 2011-01-03 | 2,612 | 2025-12-31 | 25,657 | 9.82× | 4.54% | 44.64% | +882% |
| GC | 2011-01-03 | 1,978 | 2025-12-31 | 4,351 | 2.20× | 6.28% | 13.82% | +120% |
| ZN | 2011-01-03 | 100.9 | 2025-12-31 | 112.6 | 1.12× | 1.67% | 1.86% | +12% |
| CL | 2011-01-03 | 189.6 | 2025-12-31 | 57.9 | 0.31× | 7.27% | 2.22% | −69% |
| ZC | 2011-01-03 | 723.2 | 2025-12-30 | 440.5 | 0.61× | 4.43% | 2.70% | −39% |
| NG | 2011-01-03 | 221.3 | 2025-12-31 | 3.97 | 0.018× | 7.27% | 0.13% | −98% |
Anchored pct is closest to truth near 2025-12-31; back in 2011 the stable-pct approximation under-margins high-momentum equity indices (engine accepts orders a live broker may have rejected) and over-margins crashed commodities (engine rejects orders a live broker may have accepted). The effect is confined to absolute levels in the 2011-2015 portion of the validation window; relative comparisons across families, allocators, and cost variants are unaffected. Holdout (2024–2025) is anchored at the same window as the pct calculation, so this drift does not bear on the holdout.
Account sizing: initial_cash is $10M in config/setup.yaml. The 30-product universe spans contract notionals from ≈$35k (NG) to ≈$22M (ZT 2-year T-Note), and the engine sizes positions as target_notional / contract_notional → integer contracts. At k=5 per side (10 positions total) the per-position dollar budget is 5% × cash; $10M clears the binding constraint (ES at ≈$347k holdout notional in the late window) and lets all 30 products participate. NQ in the very-late holdout (Dec 2025, peak notional ≈$513k vs $500k per-position budget) is the one residual integer-share footnote.
# From repo root
uv run python case_studies/cme_futures/01_feasibility_analysis.py
uv run python case_studies/cme_futures/02_labels.py
uv run python case_studies/cme_futures/03_financial_features.py
uv run python case_studies/cme_futures/04_model_based_features.py
uv run python case_studies/cme_futures/05_evaluation.py
uv run python case_studies/cme_futures/06_linear.py
uv run python case_studies/cme_futures/07_gbm.py
uv run python case_studies/cme_futures/08_tabular_dl.py
uv run python case_studies/cme_futures/09_dl_lstm.py
uv run python case_studies/cme_futures/10a_pca.py
uv run python case_studies/cme_futures/10b_stochastic_discount_factor.py
uv run python case_studies/cme_futures/10_latent_factors.py # summarizes 10a-10b
uv run python case_studies/cme_futures/11_causal_dml.py
uv run python case_studies/cme_futures/12_model_analysis.py
uv run python case_studies/cme_futures/13_backtest.py
uv run python case_studies/cme_futures/14_portfolio_management.py
uv run python case_studies/cme_futures/15_risk_management.py
uv run python case_studies/cme_futures/16_costs.py
uv run python case_studies/cme_futures/17_holdout_predictions.py
uv run python case_studies/cme_futures/18_holdout_backtest.py
uv run python case_studies/cme_futures/19_strategy_analysis.pyThis README describes how the case study is built, not what it found. Results are not restated here: the registry is rebuilt whenever the case study is re-derived, and a number copied into prose stays correct only until the next rebuild.
19_strategy_analysis reads the registry back and reports the selected
configuration with its interval evidence. That notebook, and the registry it reads,
are where a result comes from.
To read the results without training anything, download the published bundle, which carries the registry and the artifacts behind it:
uv run python scripts/download_artifacts.py --cs cme_futuresTwo bundles are published, and they are separate generations rather than revisions
of one another. v3.1.0-artifacts is current and is what the command above fetches.
v3.0.0-artifacts holds the results as first published. The 3.1 rebuild re-keyed
every content-addressed hash, so a hash taken from one bundle does not resolve in
the other.
run_log/registry.db records every training run, prediction set and backtest,
each addressed by a hash of the specification that produced it. The artifacts sit
beside it under run_log/training/, run_log/predictions/ and run_log/backtest/.