Skip to content

A network trained to make money, not to guess

A Deep Momentum Network, after Lim, Zohren and Roberts (University of Oxford, 2019). It learns how much to hold by maximising risk-adjusted return directly.

What we built first

An LSTM that guessed Bitcoin's 4-hour direction right 44.3% of the time, against 37.6% by chance.

It still lost money. Its edge was 0.04% a trade and the fees were 0.25%. Being right more often is not the same as making money.

What this model does instead

It is scored on profit per unit of risk, so every training step pushes toward better returns.

Bitcoin alone has too little history for a neural network, so it learns from 18 markets at once. Trends behave alike everywhere; the network learns the pattern, then applies it to Bitcoin and gold.

How one decision is made

  1. Prices

    Daily closes from 18 markets, 2000 to today

  2. 8 features

    Returns and trend signals, scaled by each market's own volatility

  3. Transformer

    Reads the last 63 trading days and outputs a position from 0 to 1

  4. Second opinion

    Averaged with eight classic trend votes on the same prices

  5. Risk

    Shrinks the position in wild markets; gold and T-bills take the rest

  6. Trade

    Decided after the US close, traded at the next open

The loss the network minimises

L = − mean(p·r) / std(p·r) × √252

p is the position the network chose, r the next day's return. Minus the annualised Sharpe ratio: the network gets better only by earning more for each unit of risk it takes.

Built in PyTorch, trained on a Kaggle GPU

Network
Transformer, 1 layer
Attention
2 heads, causal mask
Memory
63 trading days
Output
sigmoid, 0 to 1
Calibration
percentile of past 3 years
Ensemble
3 seeds averaged

It never sees the year it is tested on

Every January the network is retrained from scratch on everything before that year, then judged on the year ahead. The model was chosen on 2019 to 2023 alone; IBIT's real prices from 2024 are the exam.

2000 to 2018: learning only2019 to 2023: choosing2024 on: real exam

Every network, judged the same way

40 deep-learning variants were trained on Kaggle's GPU: 5 networks, each pooled or fine-tuned, raw or calibrated, alone or averaged with the 8 trend votes. The calibrated Transformer had the best 2019 to 2023 score on its own but earned least on real prices; averaged with the votes it was level with the rules on 2019 to 2023, kept the smallest falls on real prices, and is what runs live.

Model2019 to 2023 per yearworst fallReal ETF per yearworst fall
Transformer, calibrated+22.9%-21%+10.3%-16%
8-vote trend rules+33.1%-34%+25.9%-15%
Transformer, calibrated + 8 voteschosen+26.6%-28%+18.1%-16%
LSTM, fine-tuned + 8 votes+29.9%-38%+17.8%-20%
Buy & hold+56.5%-76%+19.9%-43%
GRU, calibrated + 8 votes+28.3%-38%+19.4%-16%
MLP, calibrated + 8 votes+26.4%-38%+19.5%-14%
TCN, fine-tuned, calibrated + 8 votes+25.6%-47%+21.8%-18%

After Indian tax at the 30% slab and all charges, in US dollars. Ranked by return per unit of worst fall on 2019 to 2023, the only years used to choose. Best of each architecture shown.

What the networks learned

Trained for risk-adjusted return, every network turned cautious: it held too little Bitcoin in strong rallies. Calibration fixed most of that; averaging with the trend votes fixed the rest.

The honest result

On real 2024 to 2026 prices the simple trend votes alone earned more. The hybrid gave up some return for a network that adapts as markets change. Both beat holding on the size of the falls.