A network trained to make money, not to guess
A Deep Momentum Network, after Lim, Zohren and Roberts (University of Oxford, 2019). It learns how much to hold by maximising risk-adjusted return directly.
What we built first
An LSTM that guessed Bitcoin's 4-hour direction right 44.3% of the time, against 37.6% by chance.
It still lost money. Its edge was 0.04% a trade and the fees were 0.25%. Being right more often is not the same as making money.
What this model does instead
It is scored on profit per unit of risk, so every training step pushes toward better returns.
Bitcoin alone has too little history for a neural network, so it learns from 18 markets at once. Trends behave alike everywhere; the network learns the pattern, then applies it to Bitcoin and gold.
How one decision is made
Prices
Daily closes from 18 markets, 2000 to today
8 features
Returns and trend signals, scaled by each market's own volatility
Transformer
Reads the last 63 trading days and outputs a position from 0 to 1
Second opinion
Averaged with eight classic trend votes on the same prices
Risk
Shrinks the position in wild markets; gold and T-bills take the rest
Trade
Decided after the US close, traded at the next open
The loss the network minimises
L = − mean(p·r) / std(p·r) × √252
p is the position the network chose, r the next day's return. Minus the annualised Sharpe ratio: the network gets better only by earning more for each unit of risk it takes.
Built in PyTorch, trained on a Kaggle GPU
- Network
- Transformer, 1 layer
- Attention
- 2 heads, causal mask
- Memory
- 63 trading days
- Output
- sigmoid, 0 to 1
- Calibration
- percentile of past 3 years
- Ensemble
- 3 seeds averaged
It never sees the year it is tested on
Every January the network is retrained from scratch on everything before that year, then judged on the year ahead. The model was chosen on 2019 to 2023 alone; IBIT's real prices from 2024 are the exam.
Every network, judged the same way
40 deep-learning variants were trained on Kaggle's GPU: 5 networks, each pooled or fine-tuned, raw or calibrated, alone or averaged with the 8 trend votes. The calibrated Transformer had the best 2019 to 2023 score on its own but earned least on real prices; averaged with the votes it was level with the rules on 2019 to 2023, kept the smallest falls on real prices, and is what runs live.
| Model | 2019 to 2023 per year | worst fall | Real ETF per year | worst fall |
|---|---|---|---|---|
| Transformer, calibrated | +22.9% | -21% | +10.3% | -16% |
| 8-vote trend rules | +33.1% | -34% | +25.9% | -15% |
| Transformer, calibrated + 8 voteschosen | +26.6% | -28% | +18.1% | -16% |
| LSTM, fine-tuned + 8 votes | +29.9% | -38% | +17.8% | -20% |
| Buy & hold | +56.5% | -76% | +19.9% | -43% |
| GRU, calibrated + 8 votes | +28.3% | -38% | +19.4% | -16% |
| MLP, calibrated + 8 votes | +26.4% | -38% | +19.5% | -14% |
| TCN, fine-tuned, calibrated + 8 votes | +25.6% | -47% | +21.8% | -18% |
After Indian tax at the 30% slab and all charges, in US dollars. Ranked by return per unit of worst fall on 2019 to 2023, the only years used to choose. Best of each architecture shown.
What the networks learned
Trained for risk-adjusted return, every network turned cautious: it held too little Bitcoin in strong rallies. Calibration fixed most of that; averaging with the trend votes fixed the rest.
The honest result
On real 2024 to 2026 prices the simple trend votes alone earned more. The hybrid gave up some return for a network that adapts as markets change. Both beat holding on the size of the falls.