Hierarchical Risk Parity Throws Away Your Return Forecast
Hierarchical Risk Parity has become a standard tool for allocating capital when covariance estimates are unreliable, yet it solves only the minimum-variance problem and cannot use an expected-return forecast at all. A recent preprint proposes three ways to close that gap, with simulation results that are strong and a limitations section that is refreshingly honest about what has not been tested.
Anyone who has built a systematic allocation process has met the same awkward moment. On one side sits a return forecast, whether it comes from a momentum rank, a valuation screen, a macro regime model or a machine-learning ensemble, and on the other side sits a covariance matrix that is too noisy to invert safely. The textbook answer, Markowitz optimisation, takes both inputs and amplifies the estimation error in the second one until the resulting weights become unusable. Hierarchical Risk Parity, proposed by López de Prado (2016), solves that instability elegantly by replacing matrix inversion with a recursive bisection over a correlation dendrogram, which is why it has spread so quickly through practitioner toolkits. What is rarely stated as plainly as it deserves is the price one pays: HRP has no place to put the return forecast. It solves the minimum-variance problem and nothing else.
A preprint posted in April 2026 by Bernd Johannes Wuebben names this property precisely, calling HRP and its Schur-complement generalisation by Cotton (2024) "signal-blind", and then constructs three allocators that keep the regularisation benefit while accepting an arbitrary expected-return vector (Wuebben, 2026). The paper runs to 93 pages and is, for the moment, an unrefereed preprint on arXiv without a journal home, which matters for how much weight one should put on it.
What signal-blindness actually costs
The diagnostic that makes the problem concrete is the out-of-sample Sharpe ratio that each method attains relative to an oracle that knows the true parameters. In Wuebben's synthetic experiments, HRP and Cotton deliver an out-of-sample Sharpe ratio near zero whenever the true expected-return vector departs from the flat vector of ones, and that flat vector is the only case for which either was built. That is not a criticism of either method, since neither was designed to consume a signal, but it does quantify what happens when a practitioner runs HRP on a universe for which a genuine forecast exists: the forecast contributes nothing to the weights, and the allocation is optimal only for an investor who believes all assets have identical expected returns.
The three proposed methods sit on a spectrum defined by a single shrinkage parameter, gamma, running from zero to one. HRP-mu keeps the recursive tree structure of the original and swaps the inverse-variance representative at each node for a signed variant that reflects the signal; it is closed-form, costs O(N squared), preserves the auditable tree, and recovers between 58 and 67 per cent of oracle Sharpe. HRP-Sigma-mu goes further by replacing the inverse-variance representative with a recursive local mean-variance optimum, thereby using within-cluster covariance information that the original discards, and it improves on HRP-mu by 20 to 35 per cent on generic random signals across sample sizes of 60, 120, 240 and 500 observations. On structurally motivated sector-tilt signals the improvement ranges more widely, from 20 to 180 per cent, with the largest gains at high gamma where the within-cluster covariance carries the most information.
CRISP, the third method, abandons the tree entirely. It solves the shrunk linear system in which the operator is a convex combination of the diagonal of the covariance matrix and the full covariance matrix, so that gamma interpolates continuously between a purely diagonal allocation rule and full Markowitz. Solved iteratively by Gauss-Seidel sweeps, it attains 80 to 94 per cent of oracle Sharpe across both signal families and all four sample sizes, and it dominates HRP, Cotton, Ledoit-Wolf shrinkage (Ledoit and Wolf, 2004) and direct Markowitz in the regimes tested. HRP-Sigma-mu reaches roughly 90 per cent of CRISP's Sharpe at pure quadratic cost with no iteration, which is the trade a practitioner who values a transparent allocation tree would probably take.
The statistical point, which is more interesting than the algorithms
Strip away the implementation detail and the paper makes one claim that deserves attention independently of whether anyone adopts these particular allocators: solve the shrunk Markowitz system rather than the exact one. The shrinkage is not a computational shortcut that trades accuracy for speed, and Wuebben is explicit that it holds whether the iterative solver happens to be faster or slower than a Cholesky factorisation. Full inversion amplifies estimation noise in the off-diagonal correlations, and damping those correlations before solving reduces the amplification more than it costs in bias. At convergence CRISP turns out to be Markowitz applied to a variance-preserving shrunk covariance, with the diagonal variances untouched and the off-diagonal correlations pulled towards zero, where the shrinkage intensity is tuned for out-of-sample Sharpe rather than for covariance-estimation loss. That last distinction is the substantive departure from the classical shrinkage literature, which optimises a matrix-norm criterion that no investor holds a position in.
Practically, the tuning story collapses to something almost embarrassingly simple. Across the sixteen noisy-signal cells of the experiment that sweeps gamma against the number of iterations, a fixed gamma of 0.5 with 100 sweeps lands within 0.02 of the global Sharpe optimum on every single cell, with a maximum gap of 0.020 and a mean of 0.009. The author also tried to derive an adaptive rule that would set gamma from observable quantities such as the ratio of sample size to universe size, the information coefficient and the noise-to-signal ratio, and he reports that the two-parameter fit achieves an out-of-sample R squared of about minus 0.11 on the validation cube, meaning it performs worse than simply predicting a constant. The directional signs match the theory, with the correlation between the empirical optimum and the sample-size ratio at plus 0.34, but the magnitudes are too weak to be prescriptive. Rather than bury that result, the paper foregrounds it and downgrades the adaptive rule to interpretive scaffolding. The reason the fit fails is itself informative: the Sharpe surface is so flat in gamma, with a median plateau width covering 38 per cent of the unit interval, that the empirical optimum wanders freely without the objective changing.
One number is offered as the thing to compute before choosing a method at all, namely the condition number of the correlation matrix. It governs both how directionally difficult the allocation problem is and how expensive the iterative solve becomes, and the paper maps four regimes of conditioning onto method and budget choices. In the worst case tested, a hedged tight-block covariance with a condition number of roughly 57,500, both equal weighting and plain HRP return what amounts to noise, and Cotton's allocator exhibits a small-sample failure mode that appears to be documented here for the first time.
What has not been tested, in the author's own words
This is where the paper distinguishes itself from most of what the literature produces, and where a reader should slow down rather than skip to the conclusion. Section 11.6 lists six limitations, and two of them are serious enough to determine how the results should be used.
The synthetic returns are Gaussian. Real asset returns are heavier-tailed, and the author states without hedging that the direction-error and Sharpe rankings could move under an elliptical or Student-t data-generating process, since no heavy-tailed simulation was run alongside the Gaussian case. Given that the entire argument concerns how estimation noise propagates through a matrix inverse, and that tail behaviour changes the noise structure of a sample covariance matrix materially, this is not a cosmetic caveat.
More decisively, there is no real-data backtest. The paper is synthetic Monte Carlo and in-sample diagnostics only, and a CRSP- or Russell-style out-of-sample study is declared out of scope and deferred. The stated reason is defensible, in that the synthetic design isolates the direction-error and Sharpe behaviour without the confounding effects of delisting, transaction costs, heavy tails and regime shifts, but the consequence stands: nothing in the paper establishes that these methods work on a live cross-section. The 80 to 94 per cent of oracle Sharpe is a statement about a simulation, not about markets.
Three further limitations are worth registering. The signal is taken as given, either as an oracle or as a sample mean, with no prior placed on it, so signal uncertainty and covariance uncertainty are never separated in any theorem, which is precisely what a Black-Litterman treatment would address. The clustering choices that determine the tree, meaning linkage type, distance metric and number of levels, affect every tree-based method in the paper, and robustness across those choices is reported only in an appendix while the main analysis uses a single default. And the comparison against Cotton is run only on the minimum-variance problem for which that allocator is defined, which is fair but means the comparison never covers the general mean-variance setting that motivates the whole paper.
A methodological finding that travels further than the paper
Buried in an appendix is a result that arguably deserves wider circulation than the allocators themselves. Portfolio weights are defined only up to positive scaling once a budget constraint is imposed, so it is natural to measure the distance between two portfolios with a scale-invariant metric such as cosine similarity on the weight vector. Wuebben shows that this convenience is one small step away from a metric that also identifies a portfolio with its own negative, and he exhibits noiseless problems on which a sum-normalised recursive mean-variance construction returns a sign-flipped portfolio, that is, the right magnitude on the wrong side of the market. A naively sign-invariant diagnostic scores those failures as near-perfect. The reported negative-cosine rate for the sum-to-one construction runs between 45 and 52 per cent, and switching to an L1 normalisation eliminates it entirely.
What makes this worth repeating is the admission attached to it. The author writes that the problem was easy to miss, that he missed it himself in early drafts, and that the bug persisted precisely because the summary metric looked healthy. The prescription he draws is general: for every scale-invariant summary metric, report a sign-sensitive companion and an out-of-sample quality metric before concluding that one method beats another. That discipline applies to any backtest reporting framework, whether or not hierarchical allocation is involved.
What this means for systematic investing
The immediate implication is narrow and checkable. If a process runs HRP while also producing a return forecast, the two are not being combined, and the allocation is answering a question nobody asked. In the RAMEP optimisation work we screen eleven allocators against each other, among them minimum variance, maximum Sharpe, Ledoit-Wolf shrinkage, equal risk contribution, inverse volatility, Black-Litterman and HRP, and the split Wuebben describes runs straight through that list: some of those methods can consume a return forecast and some structurally cannot, which makes a like-for-like ranking impossible once a forecast is in play. Identifying which of one's own allocators can actually consume a signal costs nothing and is worth doing before considering any new method.
Adopting the methods themselves is a different matter, and on current evidence premature. A 93-page preprint with strong simulation results, no peer review, no real-data validation and Gaussian-only return assumptions is a hypothesis worth testing, not a technique worth deploying. The appropriate next step is the one the author identifies as missing, namely a walk-forward study on real panels with transaction costs and realistic rebalancing, run against the signal-blind baselines on the same signal. Until that exists, the transferable content is the statistical insight about solving the shrunk system rather than the exact one, which can be tested directly within an existing pipeline by shrinking off-diagonal correlations before optimisation and measuring what happens out of sample, and the reporting discipline about pairing scale-invariant metrics with sign-sensitive ones, which costs a few lines of code and can catch a failure mode that summary statistics hide.
That is a modest takeaway from a substantial paper, which is usually how it goes when a methodological contribution has not yet met live data.
References:
Cotton, P. (2024) Schur Complementary Allocation: A Unification of Hierarchical Risk Parity and Minimum Variance Portfolios. arXiv:2411.05807 [q-fin.PM]. Available at: https://arxiv.org/abs/2411.05807 (Accessed: 12 September 2026).
Ledoit, O. and Wolf, M. (2004) 'Honey, I Shrunk the Sample Covariance Matrix', The Journal of Portfolio Management, 30(4), pp. 110-119. DOI: 10.3905/jpm.2004.110.
López de Prado, M. (2016) 'Building Diversified Portfolios that Outperform Out of Sample', The Journal of Portfolio Management, 42(4), pp. 59-69. DOI: 10.3905/jpm.2016.42.4.059.
Wuebben, B.J. (2026) Beyond De Prado and Cotton: Hierarchical and Iterative Methods for General Mean-Variance Portfolios. arXiv:2604.23833 [q-fin.PM]. Available at: https://arxiv.org/abs/2604.23833 (Accessed: 12 September 2026).
Research & Model Updates
New research, model releases, and methodology insights. No spam, no market commentary.
Double opt-in. Unsubscribe anytime. Privacy policy