Forecasting realized volatility, and what actually beats HAR
A three-term linear regression from 2009 is still the benchmark that most sophisticated alternatives fail to clear. Understanding why is more useful than any single model.
Get the target right first
GARCH infers a latent variance from daily returns. With intraday data you can simply measure it. Sum squared intraday returns over the day and you have realized variance, a consistent estimator of the day's integrated variance:
That is the good news. The bad news is that the estimator is contaminated by market microstructure noise — bid-ask bounce, discrete prices, stale quotes — and the contamination grows as sampling gets finer. Sample at one-second frequency and you are largely measuring the spread. Two standard responses:
- Sparse sampling. Five-minute returns, chosen because the noise bias has not yet dominated. Crude, but it survives contact with real data.
- Realized kernels. Barndorff-Nielsen, Hansen, Lunde and Shephard's estimator uses autocovariance corrections to stay consistent under noise, so you can use all the data instead of discarding most of it.
Separately, jumps should not be treated like diffusive variance. Bipower variation estimates the continuous part only; the difference against RV identifies the jump contribution. It matters because the two components have very different persistence — the continuous part is highly predictable, jumps essentially are not.
Before any modelling: decide the sampling frequency, the noise correction, the overnight-return treatment and the jump handling. These four choices move out-of-sample accuracy more than the choice between HAR and a neural network.
The benchmark: HAR
Corsi's (2009) Heterogeneous Autoregressive model is motivated by the idea that different participants act on different horizons — intraday traders, portfolio managers, long-term investors — and each generates volatility at their own frequency. The implementation is a regression on daily, weekly and monthly averages of past realized volatility:
It is not a long-memory model, but the cascade of three horizons reproduces hyperbolic-looking decay well enough to be indistinguishable in practice. It has no convergence problems, no likelihood surface to get lost on, and it takes milliseconds to fit.
Two refinements are almost free:
- Fit in logs. RV is right-skewed; log RV is close to Gaussian, which makes OLS far better behaved. Remember the Jensen correction when transforming back.
- Use HAR-RV-J or CHAR: separate the continuous and jump components and let the regression assign them different coefficients.
The challengers
Realized GARCH and HEAVY. Both bolt a realized measure onto a GARCH-style recursion. Realized GARCH (Hansen, Huang & Shek, 2012) adds a measurement equation linking the realized measure to latent variance, including a leverage function. HEAVY (Shephard & Sheppard) runs two equations and adapts fast after variance breaks. Both beat daily-return GARCH comfortably. Against HAR the margin is narrow and sample-dependent.
Rough volatility. Gatheral, Jaisson and Rosenbaum showed that log realized volatility behaves like fractional Brownian motion with a Hurst exponent around 0.1 — far rougher than the smooth paths standard stochastic volatility models assume. The resulting RFSV forecast formula is parsimonious and performs well. Honest caveat: the measured roughness is partly an artefact of estimation error in RV, and the debate over how much is genuine is unresolved.
Machine learning. Gradient boosting, LSTMs and temporal fusion transformers all deliver gains when the feature set is genuinely richer — cross-asset RVs, implied vol, order-flow imbalance, macro calendars. Given the same three HAR inputs, they typically match HAR and sometimes do worse. The lesson is consistent: the gain comes from the features, not the function class.
Evaluating honestly
This is where most comparisons quietly break.
- Loss function. MSE on variance is dominated by a few crisis days. QLIKE is scale-invariant and robust to noise in the volatility proxy, which is why it is the standard choice.
- Proxy error. The target is itself estimated. Patton showed only certain loss functions preserve the true ranking when the proxy is noisy; MSE and QLIKE do, many popular alternatives do not.
- Protocol. Rolling or expanding window, refit on schedule, and respect publication lags. A macro variable used on the day of release that was actually published after the close is leakage.
- Significance. A lower average loss is not a result. Use Diebold-Mariano for pairs and the Model Confidence Set when comparing many models, otherwise you are selecting on noise.
Rule of thumb: if a new model beats HAR by less than a few percent of QLIKE and does not survive a Model Confidence Set at ten percent, treat the improvement as unproven. Publish the null result anyway.
References
- Corsi, F. (2009). A Simple Approximate Long-Memory Model of Realized Volatility. Journal of Financial Econometrics 7(2), 174–196.
- Barndorff-Nielsen, O. E., Hansen, P. R., Lunde, A. & Shephard, N. (2008). Designing Realized Kernels. Econometrica 76(6), 1481–1536.
- Hansen, P. R., Huang, Z. & Shek, H. H. (2012). Realized GARCH: A Joint Model for Returns and Realized Measures of Volatility. Journal of Applied Econometrics 27(6), 877–906.
- Gatheral, J., Jaisson, T. & Rosenbaum, M. (2018). Volatility is Rough. Quantitative Finance 18(6), 933–949.
- Patton, A. J. (2011). Volatility Forecast Comparison using Imperfect Volatility Proxies. Journal of Econometrics 160(1), 246–256.
- Hansen, P. R., Lunde, A. & Nason, J. M. (2011). The Model Confidence Set. Econometrica 79(2), 453–497.
Volatility