Appendix: Deriving the Mean-Variance Frontier#
This appendix derives the formulas used in the
portfolio selection notebook, 01_markowitz. The derivation follows
Merton (1972). At the end, we check every formula numerically.
Setup#
There are \(N\) risky assets with expected returns \(\mu\) (an \(N\)-vector) and covariance matrix \(\Sigma\) (\(N \times N\)). We assume that \(\Sigma\) is positive definite, so no asset is redundant, and that the assets do not all have the same expected return. A portfolio is a weight vector \(w\) with \(w^\top \mathbf{1} = 1\). Its expected return is \(w^\top \mu\) and its variance is \(w^\top \Sigma w\).
The frontier problem#
For a target expected return \(m\),
The factor of \(\tfrac{1}{2}\) only tidies the derivative. The objective is strictly convex and the constraints are linear, so the first-order conditions identify the unique minimum.
First-order conditions#
Form the Lagrangian with multipliers \(\lambda\) and \(\gamma\):
Differentiating with respect to \(w\) and setting the result to zero,
Every frontier portfolio is a combination of the same two vectors, \(\Sigma^{-1} \mu\) and \(\Sigma^{-1} \mathbf{1}\). Only the multipliers depend on the target \(m\).
Solving for the multipliers#
Define
Multiply (1) on the left by \(\mu^\top\) and then by \(\mathbf{1}^\top\), and use the two constraints:
This is a \(2 \times 2\) linear system. Its solution is
We need \(D \neq 0\). In fact \(D > 0\): because \(\Sigma^{-1}\) is positive definite, the Cauchy-Schwarz inequality gives \(B^2 \le AC\), with equality only if \(\mu\) is proportional to \(\mathbf{1}\), which we ruled out. Substituting (2) into (1) gives the optimal weights,
The frontier is a hyperbola#
Multiply the first-order condition \(\Sigma w = \lambda \mu + \gamma \mathbf{1}\) on the left by \(w^\top\) and use the constraints again:
The variance is a parabola in \(m\). In the usual picture, with the standard deviation \(\sigma\) on the horizontal axis and \(m\) on the vertical axis, the curve is a hyperbola.
The global minimum variance portfolio#
Minimize \(\sigma^2(m)\) over \(m\): \(2Am - 2B = 0\), so \(m_{\text{GMV}} = B/A\). Then \(\lambda = 0\) and \(\gamma = 1/A\), and
With \(\lambda = 0\), the constraint on the mean does not bind, which is why \(\mu\) drops out of the weights. The portfolios on the frontier with \(m \ge B/A\) are the efficient ones. Those below have the same variance as an efficient portfolio and a lower mean.
The two-fund theorem#
Define a second portfolio, \(w_d = \Sigma^{-1} \mu / B\) (assuming \(B \ne 0\)). Then (1) can be written as
Every frontier portfolio is a mix of the same two funds. This is why the main notebook could trace the frontier by mixing two portfolios, and it is the first hint of why, in equilibrium, all investors might hold the same risky portfolio.
The tangency portfolio#
Add a risk-free asset with return \(r_f < B/A\). An investor who mixes the risk-free asset with a risky portfolio \(w\) moves along a straight line in mean-standard deviation space whose slope is the Sharpe ratio of \(w\). The best line is the steepest one:
The ratio does not change when \(w\) is scaled, so we can drop the constraint, solve, and rescale afterward. Setting the derivative of the ratio to zero gives \(\Sigma w \propto \mu - r_f \mathbf{1}\). Rescaling so that the weights sum to one,
The condition \(r_f < B/A\) says that the risk-free rate is below the expected return on the minimum variance portfolio. It guarantees that the denominator is positive and that the tangency point lies on the efficient half of the frontier.
Why the no-short problem has no closed form#
Add the constraint \(w \ge 0\), with a vector of multipliers \(\nu\). The Karush-Kuhn-Tucker conditions are
The last condition, complementary slackness, says that each asset is either held (\(w_i > 0\) and \(\nu_i = 0\)) or excluded (\(w_i = 0\)). Given the set of held assets, the conditions reduce to the unconstrained problem on that subset, and the formula above applies to it. The difficulty is finding the set. There are \(2^N - 1\) candidates, and the right one changes with \(m\).
Markowitz (1956) showed that the set changes only at a finite number of “corner portfolios” and that, between corners, the weights are linear in \(m\). His critical line algorithm walks from corner to corner. So the no-short frontier is a chain of hyperbola segments, and it is computed by an algorithm and not by a formula. We use a general-purpose numerical solver instead, since the problem is a small quadratic program.
Checking the formulas numerically#
A derivation is a claim, and claims should be tested. Using the data from the main notebook, we compare each formula with a numerical optimizer that knows nothing about the algebra.
import numpy as np
import mean_variance
import pull_crsp
df = pull_crsp.load_extract()
stocks = df.drop(columns=["MKT", "RF"])
mu, Sigma, rf = stocks.mean().values, stocks.cov().values, df["RF"].mean()
A, B, C, D = mean_variance.frontier_constants(mu, Sigma)
print(f"A = {A:.1f}, B = {B:.3f}, C = {C:.5f}, D = {D:.4f}")
assert D > 0
A = 869.2, B = 6.972, C = 0.10746, D = 44.7993
target = 0.015 # a 1.5% monthly expected return
w_formula = mean_variance.efficient_weights(mu, Sigma, target)
w_numerical = mean_variance.minimize_variance(mu, Sigma, target, allow_short=True)
print("Largest difference in weights:", np.abs(w_formula - w_numerical).max())
print("Variance from the weights: ", w_formula @ Sigma @ w_formula)
print("Variance from the formula: ", mean_variance.frontier_variance(mu, Sigma, target))
assert np.allclose(w_formula, w_numerical, atol=1e-5)
Largest difference in weights: 2.7487858006436383e-07
Variance from the weights: 0.002095557416878199
Variance from the formula: 0.002095557416878199
w_gmv_formula = np.linalg.solve(Sigma, np.ones(len(mu))) / A
w_gmv_numerical = mean_variance.minimize_variance(mu, Sigma, target=None)
print("GMV mean, B/A: ", B / A, "vs", w_gmv_numerical @ mu)
print("GMV variance, 1/A: ", 1 / A, "vs", w_gmv_numerical @ Sigma @ w_gmv_numerical)
assert np.allclose(w_gmv_formula, w_gmv_numerical, atol=1e-5)
GMV mean, B/A: 0.008020736960731644 vs 0.00802073590566311
GMV variance, 1/A: 0.0011504677635356872 vs 0.0011504677635360981
w_tan_formula = np.linalg.solve(Sigma, mu - rf) / (B - A * rf)
w_tan_numerical = mean_variance.max_sharpe_weights(mu, Sigma, rf)
print("Tangency weights sum to:", w_tan_formula.sum())
assert rf < B / A
assert np.allclose(w_tan_formula, w_tan_numerical, atol=1e-5)
Tangency weights sum to: 1.0
Finally, the claim about the no-short problem: given the set of assets that the constrained solution holds, the unconstrained formula applied to that subset reproduces the constrained weights.
w_no_short = mean_variance.minimize_variance(mu, Sigma, target, allow_short=False)
held = w_no_short > 1e-8
w_subset = mean_variance.efficient_weights(mu[held], Sigma[np.ix_(held, held)], target)
print("Assets held:", list(stocks.columns[held]))
print("Largest difference:", np.abs(w_no_short[held] - w_subset).max())
assert np.allclose(w_no_short[held], w_subset, atol=1e-5)
Assets held: ['AAPL', 'AMZN', 'NVDA', 'XOM', 'JNJ', 'PG', 'KO', 'WMT', 'CAT', 'BA']
Largest difference: 4.007085163959534e-07
References#
Markowitz, Harry. “The Optimization of a Quadratic Function Subject to Linear Constraints.” Naval Research Logistics Quarterly 3, no. 1-2 (1956): 111-133.
Merton, Robert C. “An Analytic Derivation of the Efficient Portfolio Frontier.” Journal of Financial and Quantitative Analysis 7, no. 4 (1972): 1851-1872.