How AI-assisted signals work on FX, indices and commodities
Most of what is sold as a 'trading signal' in the retail industry is a coloured arrow on a chart with no probability attached, no horizon attached, and no sizing prescription attached. That is not a signal. This is a description of what an actual signal looks like inside a working trading book, how we build them, and why the calibration step matters far more than the modelling step.
- A real signal is a conditional probability statement, scoped to an instrument and a horizon.
- Calibration matters more than ranking. A model that emits well-calibrated probabilities can be sized by edge; a model that only ranks cannot.
- Quiet models are usually better models. A signal that fires every five minutes is almost certainly overfit.
- Every signal we run is traded on our own capital before it touches any production sizing envelope.
The definition we work to
A trading signal, properly defined, is a conditional probability statement: given a specific market state, the expected return on a defined instrument over a defined horizon is positive (or negative) by a measurable margin. Everything else — the chart, the colour, the alert — is presentation.
Stripped down this way, a signal is judged on three things. First, is the probability it emits well-calibrated, meaning that signals labelled '60% likely up' do, in fact, go up roughly 60% of the time. Second, does the magnitude of the expected move justify the round-trip cost of trading it. Third, does the signal hold up out-of-sample on data the model has never touched. Most things marketed as signals fail at least one of these tests.
Features: classical and learned
Our signal pipeline blends classical quantitative features with machine-learned overlays. The classical layer captures things we know are economically meaningful: term structure of implied vs realised volatility, cross-asset correlation regimes, carry differentials in FX, inventory and seasonality features in energy, and order-flow imbalance proxies derived from publicly available tick data.
The learned layer sits on top. It is responsible for things that are harder to specify by hand: regime classification, non-linear interactions between features, and per-instrument calibration adjustments. We deliberately keep the learned layer narrow. Black-box models that consume raw price and emit positions look impressive in backtests and have, in our experience, never survived a proper walk-forward on out-of-sample data.
“A signal that is right 55% of the time on a defined horizon is genuinely valuable — but only if it is sized in proportion to its edge and its variance. Most retail signal services collapse precisely because they ignore this step.”
Calibration: the step everyone skips
A ranking model tells you which instruments are most attractive relative to each other. A calibrated model tells you the probability of each one moving in your favour. These are very different objects, and only the second one is useful for sizing.
We invest heavily in calibration. Every model output is post-processed against a held-out window using isotonic regression or a Platt-style sigmoid fit, depending on shape. We re-check calibration weekly on rolling windows. When a model goes out of calibration — typically during a regime shift — sizing is automatically reduced before the model is retrained.
The result is that we can size positions by edge and variance, not by conviction. A small but well-calibrated edge gets the size it deserves. A large-looking but poorly calibrated edge gets cut down or excluded entirely.
Coverage: where the pipeline operates
FX coverage is the G7 majors plus a curated set of EM crosses. Majors give us liquidity and clean execution; EM crosses give us a carry premium that is not fully explained by realised volatility, with the obvious caveat that EM crises are the textbook example of correlated tail events.
Indices coverage is the major US, European and Asian benchmark CFDs. Here the features lean cross-asset — momentum, volatility regime, and the term structure of implied vol. Commodities coverage is energy and metals, with explicit handling of inventory and seasonality features in the research notebooks but not always in the live signal.
Live behaviour: quiet is correct
A signal that fires constantly is almost certainly overfit, picking up noise in the data and labelling it as edge. The signal families we trust are quiet. They go weeks without firing on some instruments, fire frequently on others, and the cadence is determined by the data rather than by a marketing requirement to look busy.
Quietness is also operationally valuable. A book with five well-calibrated, quiet signals is easier to size, easier to risk-manage, and easier to attribute performance against than a book with fifty noisy ones. The simplicity tax is real, and we pay it deliberately.
Why we trade signals before sizing them
Every new signal family is traded on a small fraction of its theoretical sizing for a defined live period before being scaled. Backtests are essential, but the gap between a backtest and a live execution book is wide enough that we treat live performance as the only evidence that matters for sizing decisions.
This filter alone removes most of what is marketed in the retail signal industry. Signals that look exquisite in backtest and ordinary in live trading are the rule, not the exception. We would rather miss a year of returns on a real signal than scale a fake one into the book.
This piece is practitioner writing from a working self-trading desk. It is not investment advice. Defam AG trades only its own capital — see the disclosure page for the full statement.