<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.2.2">Jekyll</generator><link href="/feed.xml" rel="self" type="application/atom+xml" /><link href="/" rel="alternate" type="text/html" /><updated>2026-10-03T19:58:19+00:00</updated><id>/feed.xml</id><title type="html">Mike Purewal</title><subtitle>I build data driven products in Finance using Machine Learning.</subtitle><entry><title type="html">Structural Risk Models</title><link href="/2026/09/06/structural-risk-models.html" rel="alternate" type="text/html" title="Structural Risk Models" /><published>2026-09-06T00:00:00+00:00</published><updated>2026-09-06T00:00:00+00:00</updated><id>/2026/09/06/structural-risk-models</id><content type="html" xml:base="/2026/09/06/structural-risk-models.html"><![CDATA[<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Structural Risk Models</title>
<!--
  Source: github.com/meninder/portfolio-management --
  concepts/structural-risk-models/writeup.html
  Copied 2026-09-06 from commit 3ae7f79.
  GENERATED FILE -- do not edit here. The source repo is the source of truth.
  Edit writeup.html there, then run regen.py in that repo; it rewrites this
  file's .sheet block, assets/css/structural-risk.css and
  assets/js/mathjax-config.js. Those are byte-identical to the source, so
  regen.py --check is the check that they are in sync. The .pagenav and
  .preamble blocks below, and the tail of the stylesheet after its marker
  comment, are website-only chrome with no counterpart in the source.
-->
<link rel="icon" type="image/png" href="/favicon.png">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Space+Grotesk:wght@500;600;700&family=Newsreader:ital,opsz,wght@0,6..72,400;0,6..72,500;1,6..72,400&family=IBM+Plex+Mono:wght@400;500;600&display=swap" rel="stylesheet">
<link rel="stylesheet" href="/assets/css/structural-risk.css">
<script src="/assets/js/mathjax-config.js"></script>
<script defer src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/3.2.2/es5/tex-mml-chtml.js"></script><script async src="https://www.googletagmanager.com/gtag/js?id=G-9X2510WP61"></script>
<script>
  window.dataLayer = window.dataLayer || [];
  function gtag(){dataLayer.push(arguments);}
  gtag('js', new Date());

  gtag('config', 'G-9X2510WP61');
</script></head>
<body>

<nav class="pagenav"><a href="/">&larr; Mike Purewal</a> &nbsp;&middot;&nbsp; <a href="/writing/">All writing</a></nav>

<div class="preamble">
  <p class="series"><a href="/tag/grinold-kahn">Series &middot; Grinold &amp; Kahn</a></p>
  <p>I am re-reading Grinold &amp; Kahn's <em>Active Portfolio Management</em> and writing the math out as I go (using AI).  These are distillations for my own benefit.  Rest of chapters will follow.</p>
</div>

<div class="sheet">
  <h1>Structural Risk Models</h1>
  <p class="dek">Grinold &amp; Kahn, <em>Active Portfolio Management</em>, Ch.&nbsp;3 technical appendix.</p>

  <p>A covariance matrix of \(N\) assets has \(N(N+1)/2\) distinct entries.  \(N = 1000\) yields 500,500 entries.  Estimating a covariance matrix requires 83 years of monthly returns in order to have full rank.  However, an optimizer will still find a value of \(\mathbf V^{-1}\), albeit misleading. Practitioners estimate \(\mathbf V\) by imposing structure on returns rather than estimating it entry by entry.  This factor model describes returns with a handful of common drivers and assembles \(\mathbf V\) out of those instead.</p>

  <section id="setup">
    <h2><span class="num">1</span> The return decomposition</h2>

    <table class="tight">
      <tbody>
        <tr><td>\(N,\ K\)</td><td>Number of assets and number of factors, \(K \ll N\).</td></tr>
        <tr><td>\(\mathbf r\)</td><td>\(N\times 1\) asset excess returns realized over period \(t\).</td></tr>
        <tr><td>\(\mathbf X\)</td><td>\(N\times K\) <strong>exposure matrix</strong>. \(X_{nk}\) is asset \(n\)'s exposure to factor \(k\). Known at the end of \(t-1\).</td></tr>
        <tr><td>\(\mathbf b\)</td><td>\(K\times 1\) <strong>factor returns</strong> realized over \(t\). Not directly observed.</td></tr>
        <tr><td>\(\mathbf u\)</td><td>\(N\times 1\) <strong>specific returns</strong> &mdash; the part of \(\mathbf r\) no factor explains.</td></tr>
        <tr><td>\(\mathbf F\)</td><td>\(K\times K\) covariance matrix of \(\mathbf b\).</td></tr>
        <tr><td>\(\boldsymbol\Delta\)</td><td>\(N\times N\) covariance matrix of \(\mathbf u\); assumed <strong>diagonal</strong>, \(\Delta_{nn} = \operatorname{Var}(u_n)\).</td></tr>
        <tr><td>\(\mathbf h,\ \mathbf x\)</td><td>\(N\times 1\) holdings and the portfolio's \(K\times 1\) factor exposures, \(\mathbf x = \mathbf X^{T}\mathbf h\).</td></tr>
      </tbody>
    </table>

    <p>The model is built in four steps:</p>

    <ol class="plain">
      <li>Section (&#167;2.1). Read today's exposures \(\mathbf X_t\) off company data &mdash; no returns involved.</li>
      <li>Section (&#167;2.2). For each past period \(t\), regress that period's realized returns for each stock (\(\mathbf r_t\)) on the exposures (\(\mathbf X_{t-1}\)) known before it. The coefficients are the factor returns \(\hat{\mathbf b}\), one set per period \(t\). The section then shows how the history of \(\hat{\mathbf b}\) gives \(\mathbf F\), and the leftover residuals give \(\boldsymbol\Delta\). The same section shows that this estimator is itself a set of portfolios, one per factor.</li>
      <li>Section (&#167;2.3). Steps 1 and 2 produce the three objects needed to assemble the covariance matrix: \(\hat{\mathbf V}_{t+1} = \mathbf X_t\mathbf F\mathbf X_t^{T} + \boldsymbol\Delta\).</li>
      <li>Section (&#167;3). Split a portfolio's risk using this.</li>
    </ol>
    <div class="def">
      <div class="label">The decomposition</div>
      <p style="margin:0">Every asset's return over period \(t\) splits into a part explained by its factor exposures and a residual. The exposures are dated \(t-1\) because they are fixed before the period runs:</p>
      \[ \mathbf r_t = \mathbf X_{t-1}\,\mathbf b_t + \mathbf u_t \tag{1} \]
    </div>

    <p class="note">The period subscripts are dropped below wherever only one period is in play.</p>
  </section>

  <section id="exposures">
    <p class="eyebrow">Part 2 &mdash; Building the model</p>
    <h2><span class="num">2.1</span> Building the Exposure Matrix \(\mathbf X\)</h2>

    <p>Every entry of \(\mathbf X_{t-1}\) has to be measurable before period \(t\) begins, which rules out anything derived from the returns it is meant to explain. The columns come in two kinds.</p>

    <p><b>Industry columns.</b> \(X_{nk} = 1\) if asset \(n\) is in industry \(k\), else 0; conglomerates get fractional membership by sales or assets, and the row still sums to one. These absorb the co-movement that comes from doing the same business.</p>

    <p><b>Risk-index columns.</b> Continuous descriptors &mdash; size, value, momentum, volatility, leverage, liquidity, growth. Raw descriptors are in incompatible units (log market cap against a book-to-price ratio), so each is standardized across the estimation universe:</p>

    \[ X_{nk} = \frac{d_{nk} - \mu_k}{\sigma_k} \tag{2} \]

    <p>The Barra convention takes \(\mu_k\) <strong>cap-weighted</strong> and \(\sigma_k\) <strong>equal-weighted</strong>. The two choices do different jobs, and the cap-weighted \(\mu_k\) is doing something exact. Let \(w_n\) be the benchmark's cap weights, so \(\sum_n w_n = 1\), and let \(\mu_k\) be defined with those same weights:</p>

    \[ \mu_k \;\equiv\; \sum_n w_n d_{nk} \]

    <p>Then the benchmark's own exposure to risk index \(k\) is</p>

    \begin{align}
      \sum_n w_n X_{nk} &= \sum_n w_n\,\frac{d_{nk} - \mu_k}{\sigma_k} && \text{by eq.\,2} \\[4pt]
      &= \frac{1}{\sigma_k}\Big(\sum_n w_n d_{nk} - \mu_k \sum_n w_n\Big) && \sigma_k,\ \mu_k\text{ do not depend on }n \\[4pt]
      &= \frac{1}{\sigma_k}\Big(\sum_n w_n d_{nk} - \mu_k\Big) && \sum_n w_n = 1 \\[4pt]
      &= 0 && \text{by the definition of }\mu_k \tag{3}
    \end{align}

    <p>So it is true by construction, and only for the one portfolio whose weights were used to centre the descriptor. An equal-weighted \(\mu_k\) would leave a non-zero benchmark exposure, and so would cap weights from a different index.</p>

    <p>That zero is what makes the exposures readable: a portfolio's exposure is directly a bet against the benchmark, and \(\mathbf x = \mathbf X^{T}\mathbf h\) needs no re-centering before it is interpreted. The equal-weighted \(\sigma_k\) keeps a handful of megacaps from setting the scale, so a one-unit exposure means one cross-sectional standard deviation across the universe rather than across the index's largest members.</p>

    <p><b>Four practical points</b>, none of which appear in the algebra but all of which decide whether the fitted model is any good:</p>

    <ul class="plain caveats">
      <li><b>Winsorization.</b> The cross-sectional regression is least squares, so descriptors are standardized, trimmed at roughly \(\pm 3\) standard deviations, then standardized again.</li>
      <li><b>Point-in-time data.</b> Exposures must use what was knowable at \(t-1\): as-first-reported fundamentals, not restatements, and filing dates rather than fiscal period ends. Otherwise \(\mathbf X\) is not genuinely known before the period and the backtest inherits look-ahead.</li>
      <li><b>Composite descriptors.</b> A <em>descriptor</em> is one raw measured quantity; a risk-index column is usually several of them standardized separately and then averaged. Value combines book-to-price with earnings-to-price and cash-flow-to-price; momentum combines returns over several horizons. Each is a noisy proxy for the same underlying quality &mdash; book-to-price is distorted by intangibles, earnings-to-price by one-off items &mdash; so averaging cancels the idiosyncratic measurement error and keeps what they share.</li>
      <li><b>Missing data.</b> Assign the peer-group mean rather than dropping the asset, which after standardization is an exposure at or near zero. Dropping assets shrinks the cross-section and biases it toward larger, better-covered names.</li>
    </ul>
  </section>

  <section id="estimation">
    <h2><span class="num">2.2</span> Estimating the Factor Returns \(\mathbf b\)</h2>

    <p>Fix a period \(t\) and write eq. 1 out one asset at a time, with \(X_{nk}\) the \((n,k)\) entry of \(\mathbf X_{t-1}\):</p>

    \[ r_{nt} \;=\; \sum_{k=1}^{K} X_{nk}\,b_{kt} \;+\; u_{nt}, \qquad n = 1,\dots,N \tag{4} \]

    <p>That is \(N\) equations sharing \(K\) unknowns \(b_{1t},\dots,b_{Kt}\). Every quantity in eq. 4 belongs to the one period \(t\), and the index that runs is \(n\): <strong>the observations are assets, not dates</strong>. Since \(K \ll N\) the system is overdetermined, so it is solved by least squares, and the whole thing is run again from scratch next period.</p>

    <p>Getting that direction backwards is the most common misreading of the model, because the better-known construction runs the other way:</p>

    <table class="tight">
      <tbody>
        <tr><td></td><td><b>Here</b></td><td><b>Fama&ndash;French</b></td></tr>
        <tr><td>Exposures</td><td>known, from company characteristics</td><td>estimated</td></tr>
        <tr><td>Factor returns</td><td>estimated</td><td>known, as constructed portfolio returns</td></tr>
        <tr><td>A regression spans</td><td>the \(N\) stocks in one period</td><td>the \(T\) months for one stock</td></tr>
        <tr><td>Regressions run</td><td>one per period</td><td>one per stock</td></tr>
        <tr><td>Exposures change</td><td>when the company does</td><td>when a rolling window catches up</td></tr>
        <tr><td>A stock with no history</td><td>still has a cap and an industry, so it has exposures</td><td>cannot be fitted</td></tr>
      </tbody>
    </table>

    <p><b>OLS to get \(\mathbf b\).</b> Minimizing \(\mathbf u^{T}\mathbf u = (\mathbf r - \mathbf X\mathbf b)^{T}(\mathbf r - \mathbf X\mathbf b)\) gives</p>

    \[ \hat{\mathbf b} = \left(\mathbf X^{T}\mathbf X\right)^{-1}\mathbf X^{T}\mathbf r \tag{5} \]

    <p>which weights every stock equally. But specific risk is strongly heteroskedastic &mdash; a micro-cap's specific volatility runs several times a mega-cap's &mdash; so equal weighting lets the noisiest assets in the cross-section drive the estimate. Weighting each observation by its precision \(1/\Delta_{nn}\) instead means minimizing \((\mathbf r - \mathbf X\mathbf b)^{T}\boldsymbol\Delta^{-1}(\mathbf r - \mathbf X\mathbf b)\):</p>

    <div class="keybox">
      <div class="cap">Cross-sectional GLS</div>
      \[ \hat{\mathbf b} = \left(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\right)^{-1}\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf r \tag{6} \]
    </div>

    <p>By Gauss&ndash;Markov this is the minimum-variance linear unbiased estimator when the residuals are uncorrelated with unequal variances, which is precisely a diagonal \(\boldsymbol\Delta\) with unequal entries &mdash; though only with the true \(\boldsymbol\Delta\), since substituting an estimate for it matches the bound only asymptotically. OLS is the special case \(\boldsymbol\Delta = \sigma^{2}\mathbf I\): equal weighting is optimal only if every stock carries the same specific risk, which no equity universe does. In practice \(1/\Delta_{nn}\) is often approximated by market cap or its square root, since specific variance falls roughly with size.</p>

    <p><b>Full column rank.</b> Both estimators invert a \(K\times K\) matrix built from \(\mathbf X\), so \(\mathbf X\) must have rank \(K\) or neither has a unique solution. The industry columns of &#167;2.1 are the standard way this fails: every asset belongs somewhere, so those columns sum to \(\mathbf e\), the vector of ones, and adding a market or intercept column makes \(\mathbf X\) rank-deficient. Two standard fixes: drop one industry, which makes every industry return relative to the omitted one, or keep all of them and impose a constraint such as \(\sum_k c_k b_k = 0\), where \(c_k\) is industry \(k\)'s share of total market capitalization, which keeps the industries symmetric and interprets the intercept as the market return.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:28px 0 10px">The estimator is a set of portfolios</h3>

    <p>Equation 6 is linear in \(\mathbf r\). Write it as \(\hat{\mathbf b} = \mathbf H^{T}\mathbf r\) and read off the \(N\times K\) matrix that does it:</p>

    \[ \mathbf H^{T} = \left(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\right)^{-1}\mathbf X^{T}\boldsymbol\Delta^{-1}, \qquad \mathbf H = \boldsymbol\Delta^{-1}\mathbf X\left(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\right)^{-1} \tag{7} \]

    <p>Column \(k\) of \(\mathbf H\) is a vector of holdings, and \(\hat b_k = \mathbf h_k^{T}\mathbf r\) is that portfolio's return. So <strong>each estimated factor return is the realized return on an actual portfolio</strong>, not a statistical abstraction. Whatever the factor did this month, some book of long and short positions earned exactly that.</p>

    <p>What those portfolios hold is pinned down by one identity:</p>

    \[ \mathbf H^{T}\mathbf X = \left(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\right)^{-1}\left(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\right) = \mathbf I_K \tag{8} \]

    <p>Entry \((k,l)\) of \(\mathbf H^{T}\mathbf X\) is portfolio \(k\)'s exposure to factor \(l\). So portfolio \(k\) has <em>unit</em> exposure to factor \(k\) and <em>zero</em> exposure to every other factor: it is a pure bet on one factor, hedged against the rest.</p>

    <p><b>What these portfolios minimize.</b> For a single attribute \(\mathbf a\), the minimum-variance portfolio holding one unit of exposure to it &mdash; \(\min_{\mathbf h}\mathbf h^{T}\mathbf V\mathbf h\) subject to \(\mathbf h^{T}\mathbf a = 1\) &mdash; is called the <em>characteristic portfolio</em> of \(\mathbf a\), and works out to \(\mathbf V^{-1}\mathbf a/(\mathbf a^{T}\mathbf V^{-1}\mathbf a)\). Column \(k\) of eq. 7 solves the same problem with \(\boldsymbol\Delta\) in place of \(\mathbf V\), and \(K\) constraints in place of one:</p>

    \[ \min_{\mathbf h}\ \mathbf h^{T}\boldsymbol\Delta\mathbf h \qquad\text{s.t.}\qquad \mathbf X^{T}\mathbf h = \mathbf e_k \tag{9} \]

    <p><a href="#lagrange">Appendix&nbsp;A</a> does the Lagrangian; at \(K = 1\) it collapses to \(\boldsymbol\Delta^{-1}\mathbf a/(\mathbf a^{T}\boldsymbol\Delta^{-1}\mathbf a)\), the characteristic-portfolio formula with \(\mathbf V \to \boldsymbol\Delta\).</p>

    <p class="note"><b>They are estimation devices, not trades.</b> Nothing in eq. 9 mentions turnover, borrow, position limits, or transaction costs. Factor portfolios typically hold thousands of names at tiny weights, carry large gross exposure relative to net, rebalance completely every period by construction, and take positions in assets no one can short. Tradeable factor products are separate objects, built with those constraints in the optimization.</p>
    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:28px 0 10px">From the estimates to \(\mathbf F\) and \(\boldsymbol\Delta\)</h3>

    <p><b>The circularity.</b> Equation 6 needs \(\boldsymbol\Delta\), but \(\boldsymbol\Delta\) is estimated from the variance of the regression residuals \(\hat{\mathbf u} = \mathbf r - \mathbf X\hat{\mathbf b}\), which needs \(\hat{\mathbf b}\). Two ways out, usually combined: iterate within a period &mdash; start from OLS or a size proxy, form residuals, re-weight, repeat until the estimates settle &mdash; or use \(\hat{\boldsymbol\Delta}\) from prior periods, which is what production systems do, since a lagged estimate cannot introduce look-ahead and specific variances move slowly.</p>

    <p><b>From the panel to \(\mathbf F\) and \(\boldsymbol\Delta\).</b> Running eq. 6 each month for \(t = 1,\dots,T\) gives a panel \(\{\hat{\mathbf b}_t\}\) of \(K\) time series. Then \(\mathbf F\) is the sample covariance matrix of those series and \(\Delta_{nn}\) is the time-series variance of asset \(n\)'s residuals \(\hat u_{nt}\). Both are ordinary time-series estimates &mdash; but now on \(K\) series rather than \(N\), which is the whole reason the exercise is tractable.</p>

    <p>Each period's regression used a different \(\mathbf X_{t-1}\), so the panel is only comparable across periods because the descriptors are re-standardized every period by eq. 2. That is what fixes the units: \(\hat b_{tk}\) means the same thing in every \(t\) &mdash; the return to a one-cross-sectional-standard-deviation bet on characteristic \(k\) &mdash; so covariances taken down the panel are meaningful. Standardization is not cosmetic; without it factor returns from different decades would be denominated differently and \(\mathbf F\) would be an average of incompatible numbers.</p>

    <p class="note">\(\hat{\mathbf b}_t\) is an estimate, not the factor return itself, so its sample covariance carries estimation error on top of the true \(\mathbf F\) and is biased upward. Production models correct for this, and typically also shrink \(\mathbf F\) and scale it to match realized volatility.</p>
  </section>

  <section id="covariance">
    <h2><span class="num">2.3</span> The covariance result</h2>

    <p>Take \(\mathbf b\) and \(\mathbf u\) centered. \(\mathbf V\) is a covariance, which is mean-insensitive. Then \(\mathbf V = \operatorname{Var}(\mathbf r) = \mathbb E[\mathbf r\mathbf r^{T}]\).</p>

    \begin{align}
      \mathbf V &= \mathbb E\big[(\mathbf X\mathbf b + \mathbf u)(\mathbf X\mathbf b + \mathbf u)^{T}\big] && \text{by eq. 1} \\[4pt]
      &= \mathbb E\big[\mathbf X\mathbf b\mathbf b^{T}\mathbf X^{T}\big] + \mathbb E\big[\mathbf X\mathbf b\mathbf u^{T}\big] + \mathbb E\big[\mathbf u\mathbf b^{T}\mathbf X^{T}\big] + \mathbb E\big[\mathbf u\mathbf u^{T}\big] && \text{expand} \\[4pt]
      &= \mathbf X\,\mathbb E[\mathbf b\mathbf b^{T}]\,\mathbf X^{T} + \mathbf X\,\mathbb E[\mathbf b\mathbf u^{T}] + \mathbb E[\mathbf u\mathbf b^{T}]\,\mathbf X^{T} + \mathbb E[\mathbf u\mathbf u^{T}] && \mathbf X\text{ is known at }t-1 \\[4pt]
      &= \mathbf X\mathbf F\mathbf X^{T} + \mathbf X\operatorname{Cov}(\mathbf b,\mathbf u) + \operatorname{Cov}(\mathbf u,\mathbf b)\,\mathbf X^{T} + \boldsymbol\Delta && \text{definitions of }\mathbf F,\ \boldsymbol\Delta \\[4pt]
      &= \mathbf X\mathbf F\mathbf X^{T} + \boldsymbol\Delta && \operatorname{Cov}(\mathbf b,\mathbf u) = \mathbf 0 \tag{10}
    \end{align}

    <p><b>Equation 10 read as a forecast.</b> Standing at date \(t\) and wanting next month's covariance matrix, eq. 10 says</p>

    \[ \hat{\mathbf V}_{t+1} = \mathbf X_t\,\hat{\mathbf F}_t\,\mathbf X_t^{T} + \hat{\boldsymbol\Delta}_t \tag{11} \]

    <p>\(\mathbf X_t\) determines the exposures that will govern next month's returns based on today's market capitalizations, book-to-price ratios and industry memberships. \(\mathbf F\) and \(\boldsymbol\Delta\) are estimated from history, so they describe factor behaviour in general. \(\hat{\mathbf F}_t\) is \(K\times K\), and the same \(\hat{\mathbf F}_t\) is used for every pair of stocks, so it cannot be what makes one pair differ from another. \(\mathbf X_t\) is \(N\times K\), one row per stock: row \(n\) holds stock \(n\)'s own characteristics, so Apple's row and Pfizer's row are different numbers, and that difference is the only thing separating one covariance from another.</p>

    <p>A sample covariance matrix has one input and one only, past returns, with no slot for the fact that a company is now ten times the size it was. Five years a micro-cap and then a tripling leaves sixty monthly observations of which fifty-nine describe the small version of it, and they clear the window one month at a time. Recomputing \(\mathbf X_t\) today moves that stock's size exposure at once. The exception is \(\hat{\boldsymbol\Delta}_t\), one variance per asset taken from that asset's own past residuals, so specific risk lags in exactly the way a sample estimate does.</p>

    <p>The subscripts on \(\hat{\mathbf F}_t\) and \(\hat{\boldsymbol\Delta}_t\) mark them as estimated from history through \(t\), and that history was built from the <em>past</em> exposures \(\mathbf X_0,\dots,\mathbf X_{t-1}\) (&#167;2.2). Today's \(\mathbf X_t\) appears in eq. 11 and nowhere else &mdash; it is information no historical estimate contains.</p>

    <div class="def warn">
      <div class="label">Assumptions used</div>
      <p style="margin:0 0 8px"><b>Factor and specific returns are uncorrelated:</b> \(\operatorname{Cov}(\mathbf b,\mathbf u) = \mathbf 0\). This is what kills the two cross terms in the derivation above.</p>
      <p style="margin:0"><b>\(\boldsymbol\Delta\) is diagonal:</b> specific returns are mutually uncorrelated. Note this one is <em>not</em> needed for eq. 10 &mdash; the identity holds for any \(\boldsymbol\Delta\). It is needed for the entry-by-entry reading and the parameter count below.</p>
    </div>

    <p><b>Dimension check.</b> \(\mathbf X\) is \(N\times K\), \(\mathbf F\) is \(K\times K\), \(\mathbf X^{T}\) is \(K\times N\), so \(\mathbf X\mathbf F\mathbf X^{T}\) is \(N\times N\) and conformable with \(\boldsymbol\Delta\). Both terms are symmetric, so \(\mathbf V\) is.</p>

    <p><b>Entry by entry.</b></p>

    \[ V_{ij} = \sum_{k}\sum_{l} X_{ik}F_{kl}X_{jl} + \Delta_{ij} \tag{12} \]

    <p>Because \(\boldsymbol\Delta\) is diagonal the second term vanishes for \(i \ne j\), so <strong>two different assets covary only through their factor exposures</strong>. Two biotech firms are correlated because both load on the same industry and size factors, not because the model was told anything about the pair. The converse is the restriction: two stocks with identical rows of \(\mathbf X\) get identical covariance with every other asset, and differ only in \(\Delta_{nn}\).</p>

    <p><b>Parameter count.</b> \(\mathbf F\) contributes \(K(K+1)/2\) and diagonal \(\boldsymbol\Delta\) contributes \(N\). Sixty factors over a thousand assets is \(1830 + 1000 = 2830\) parameters against 500,500 &mdash; 177 times fewer. The \(NK\) numbers in \(\mathbf X\) are not estimated at all.</p>

  </section>

  <section id="using">
    <p class="eyebrow">Part 3 &mdash; Using the model</p>
    <h2><span class="num">3</span> Risk Decomposition</h2>

    <p>Substituting eq. 10 into portfolio variance and collecting \(\mathbf X^{T}\mathbf h\):</p>

    \begin{align}
      \sigma_P^{2} = \mathbf h^{T}\mathbf V\mathbf h &= \mathbf h^{T}\left(\mathbf X\mathbf F\mathbf X^{T} + \boldsymbol\Delta\right)\mathbf h && \text{by eq. 10} \\[4pt]
      &= \left(\mathbf X^{T}\mathbf h\right)^{T}\mathbf F\left(\mathbf X^{T}\mathbf h\right) + \mathbf h^{T}\boldsymbol\Delta\mathbf h && \text{regroup} \\[4pt]
      &= \underbrace{\mathbf x^{T}\mathbf F\mathbf x}_{\text{common factor}} + \underbrace{\mathbf h^{T}\boldsymbol\Delta\mathbf h}_{\text{specific}} \tag{13}
    \end{align}

    <p>Every attribution report is a reading of eq. 13. The first term depends on the portfolio only through its \(K\) factor exposures \(\mathbf x\), so two portfolios with different holdings and the same exposures carry identical common-factor risk. The second is \(\sum_n h_n^{2}\Delta_{nn}\) for diagonal \(\boldsymbol\Delta\), which shrinks like \(1/N\) as holdings spread out: an equal-weighted book drives specific risk toward zero while its common-factor risk stays put. That is the precise sense in which factor risk is the part diversification cannot remove.</p>

    <p><b>Why &#167;2.2 could minimize over \(\boldsymbol\Delta\).</b> Using \(\boldsymbol\Delta\) rather than \(\mathbf V\) changes nothing, because the constraint has already spent the factor risk:</p>

    \begin{align}
      \mathbf h^{T}\mathbf V\mathbf h &= \mathbf x^{T}\mathbf F\mathbf x + \mathbf h^{T}\boldsymbol\Delta\mathbf h && \text{by eq. 13} \\[4pt]
      &= \mathbf e_k^{T}\mathbf F\mathbf e_k + \mathbf h^{T}\boldsymbol\Delta\mathbf h && \mathbf x = \mathbf X^{T}\mathbf h = \mathbf e_k\text{, by eq. 9} \\[4pt]
      &= F_{kk} + \mathbf h^{T}\boldsymbol\Delta\mathbf h \tag{14}
    \end{align}

    <p>\(F_{kk}\) is the same number for every feasible \(\mathbf h\), so minimizing total variance and minimizing specific variance are the same problem here and have the same minimizer. \(\boldsymbol\Delta\) appears because GLS weights by the covariance of the regression's errors, and it survives into the portfolio statement because requiring unit exposure to one factor and zero to the rest has already fixed the common-factor variance at \(F_{kk}\).</p>
  </section>

  <section id="assumptions">
    <h2><span class="num">4</span> Assumptions</h2>

    <p>What the construction leans on.</p>

    <ul class="plain caveats">
      <li><b>\(\mathbf X\) is known at \(t-1\) and fixed through the period.</b> Used to pull \(\mathbf X\) out of the expectation in eq. 10. Exposures drift within a month as prices move &mdash; a stock's size and momentum exposures at month-end are not the ones the model used.</li>
      <li><b>\(\operatorname{Cov}(\mathbf b,\mathbf u) = \mathbf 0\).</b> Not directly testable, since both sides are estimated from the same regression, and the residuals are orthogonal to \(\mathbf X\) by construction whether or not the population statement holds. A genuinely omitted factor that loads on the included characteristics violates it.</li>
      <li><b>\(\boldsymbol\Delta\) diagonal.</b> The strongest assumption in the model. Real residual correlation survives any factor set &mdash; dual listings, merger pairs, supply-chain partners, companies sharing a single customer. Concentrated portfolios are where the understatement bites, since eq. 13 then treats correlated bets as independent ones.</li>
      <li><b>Linearity.</b> Equation 1 admits no interaction or nonlinear terms; an effect that appears only for small <em>and</em> illiquid names has nowhere to go but \(\mathbf u\).</li>
      <li><b>Stationarity of \(\mathbf F\) and \(\boldsymbol\Delta\)</b> over the estimation window, so that a sample covariance of past \(\hat{\mathbf b}_t\) forecasts the next period. Factor volatilities are visibly regime-dependent, which is why production models rescale.</li>
      <li><b>\(\boldsymbol\Delta \succ 0\).</b> Needed for \(\mathbf V \succ 0\) and for the GLS weighting in eq. 6. \(\mathbf X\mathbf F\mathbf X^{T}\) alone has rank at most \(K\).</li>
      <li><b>\(\mathbf X\) has full column rank \(K\).</b> Needed for eqs. 5 and 6 to have unique solutions; the industry-plus-intercept collinearity of &#167;2.1 is the standard way this fails.</li>
      <li><b>It is a risk model, not a return model.</b> Nothing here estimates \(\mathbb E[\mathbf b]\), and the factor structure carries no claim that exposure is compensated. Using \(\mathbf X\mathbb E[\mathbf b]\) as an alpha is a separate assertion needing separate evidence.</li>
    </ul>
  </section>

  <section id="lagrange">
    <h2><span class="num">A</span> Appendix &mdash; factor portfolios as a constrained minimization</h2>

    <p>Claim: column \(k\) of \(\mathbf H\) in eq. 7 solves eq. 9. The Lagrangian carries a \(K\times 1\) multiplier \(\boldsymbol\lambda\), one per constraint:</p>

    \begin{align}
      L(\mathbf h,\boldsymbol\lambda) &= \tfrac12\,\mathbf h^{T}\boldsymbol\Delta\mathbf h - \boldsymbol\lambda^{T}\left(\mathbf X^{T}\mathbf h - \mathbf e_k\right) && \text{one multiplier per constraint} \tag{A1} \\[4pt]
      \boldsymbol\Delta\mathbf h - \mathbf X\boldsymbol\lambda &= \mathbf 0 && \partial L/\partial\mathbf h = \mathbf 0 \\[4pt]
      \mathbf h &= \boldsymbol\Delta^{-1}\mathbf X\boldsymbol\lambda && \boldsymbol\Delta \succ 0 \tag{A2}
    \end{align}

    <p>Imposing the constraint on eq. A2 determines \(\boldsymbol\lambda\), and substituting back gives \(\mathbf h\):</p>

    \begin{align}
      \mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\boldsymbol\lambda &= \mathbf e_k && \text{apply }\mathbf X^{T}\text{ to eq. A2} \\[4pt]
      \boldsymbol\lambda &= \left(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\right)^{-1}\mathbf e_k && \mathbf X\text{ full rank} \\[4pt]
      \mathbf h_k &= \boldsymbol\Delta^{-1}\mathbf X\left(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X\right)^{-1}\mathbf e_k = \mathbf H\mathbf e_k \tag{A3}
    \end{align}

    <p>which is column \(k\) of eq. 7. The objective is strictly convex and the constraints affine, so this is the unique minimizer, not merely a stationary point.</p>

    <p><b>The \(K = 1\) case.</b> With one factor, \(\mathbf X\) is a single column \(\mathbf a\) and \(\mathbf e_k = 1\). Then \(\mathbf X^{T}\boldsymbol\Delta^{-1}\mathbf X = \mathbf a^{T}\boldsymbol\Delta^{-1}\mathbf a\) is a scalar and eq. A3 reads</p>

    \[ \mathbf h_a = \frac{\boldsymbol\Delta^{-1}\mathbf a}{\mathbf a^{T}\boldsymbol\Delta^{-1}\mathbf a} \tag{A4} \]

    <p>The characteristic portfolio of \(\mathbf a\) from <a href="#estimation">&#167;2.2</a>, with \(\boldsymbol\Delta\) wherever \(\mathbf V\) stood. The \(K &gt; 1\) version adds the requirement that the portfolio be neutral to the other \(K-1\) factors.</p>
  </section>

</div>

<nav class="pagenav foot"><a href="/">&larr; Mike Purewal</a> &nbsp;&middot;&nbsp; <a href="/writing/">All writing</a> &nbsp;&middot;&nbsp; <a href="/tag/grinold-kahn">More in this series</a></nav>

</body>
</html>]]></content><author><name></name></author><category term="grinold-kahn" /><category term="finance" /><category term="mathematics" /><summary type="html"><![CDATA[Structural Risk Models]]></summary></entry><entry><title type="html">Characteristic Portfolios</title><link href="/2026/08/30/characteristic-portfolios.html" rel="alternate" type="text/html" title="Characteristic Portfolios" /><published>2026-08-30T00:00:00+00:00</published><updated>2026-08-30T00:00:00+00:00</updated><id>/2026/08/30/characteristic-portfolios</id><content type="html" xml:base="/2026/08/30/characteristic-portfolios.html"><![CDATA[<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Characteristic Portfolios</title>
<!--
  Source: github.com/meninder/portfolio-management --
  concepts/characteristic-portfolios/writeup.html
  Copied 2026-08-30 from commit 19f5322.
  GENERATED FILE -- do not edit here. The source repo is the source of truth.
  Edit writeup.html there, then regenerate this file's .sheet block,
  assets/css/char-portfolios.css and assets/js/mathjax-config.js.
  All three are byte-identical to the source, so diff is the check that they
  are in sync. The .pagenav and .preamble blocks below, and the tail of the
  stylesheet after its marker comment, are website-only chrome with no
  counterpart in the source.
-->
<link rel="icon" type="image/png" href="/favicon.png">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Space+Grotesk:wght@500;600;700&family=Newsreader:ital,opsz,wght@0,6..72,400;0,6..72,500;1,6..72,400&family=IBM+Plex+Mono:wght@400;500;600&display=swap" rel="stylesheet">
<link rel="stylesheet" href="/assets/css/char-portfolios.css">
<script src="/assets/js/mathjax-config.js"></script>
<script defer src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/3.2.2/es5/tex-mml-chtml.js"></script><script async src="https://www.googletagmanager.com/gtag/js?id=G-9X2510WP61"></script>
<script>
  window.dataLayer = window.dataLayer || [];
  function gtag(){dataLayer.push(arguments);}
  gtag('js', new Date());

  gtag('config', 'G-9X2510WP61');
</script></head>
<body>

<nav class="pagenav"><a href="/">&larr; Mike Purewal</a> &nbsp;&middot;&nbsp; <a href="/writing/">All writing</a></nav>

<div class="preamble">
  <p class="series"><a href="/tag/grinold-kahn">Series &middot; Grinold &amp; Kahn</a></p>
  <p>I am re-reading Grinold &amp; Kahn's <em>Active Portfolio Management</em> and writing the math out as I go (using AI).  These are distillations for my own benefit.  Rest of chapters will follow.</p>
</div>

<div class="sheet">

  
  <h1>Characteristic Portfolios</h1>
  <p class="dek">Grinold &amp; Kahn, <em>Active Portfolio Management</em>, Ch.&nbsp;2 technical appendix.</p>

  <section id="setup">
    <h2><span class="num">01</span> Setup</h2>

        <table class="tight">
      <tbody>
        <tr><td>\(N\)</td><td>Number of risky assets.</td></tr>
        <tr><td>\(\mathbf r\)</td><td>\(N\times 1\) vector of asset excess returns, \(r_n\) the excess return on asset \(n\).</td></tr>
        <tr><td>\(\mathbf V\)</td><td>\(N\times N\) covariance matrix of those returns. Symmetric, <strong>positive definite</strong>; \(\mathbf V^{-1}\) exists, symmetric.</td></tr>
        <tr><td>\(\mathbf a\)</td><td>\(N\times 1\) vector of asset attributes. Anything measurable per asset: book-to-price, a beta estimate, a vector of ones.</td></tr>
        <tr><td>\(\mathbf h\)</td><td>\(N\times 1\) holdings, \(h_n\) the fraction of portfolio value in asset \(n\).</td></tr>
        <tr><td>\(\mathbf e\)</td><td>\(N\times 1\) vector of ones, \((1,\dots,1)^{T}\). So \(\mathbf h^{T}\mathbf e\) is total holdings, and \(\mathbf h^{T}\mathbf e = 1\) says the portfolio is fully invested.</td></tr>
      </tbody>
    </table>

    \begin{align}
      \text{Portfolio excess return:}\quad & r_P = \mathbf h_P^{T}\mathbf r \tag{1.1} \\[6pt]
      \text{Portfolio variance:}\quad & \sigma_P^{2} = \operatorname{Var}\!\left(\mathbf h_P^{T}\mathbf r\right) = \mathbf h_P^{T}\operatorname{Var}(\mathbf r)\,\mathbf h_P = \mathbf h_P^{T}\mathbf V\mathbf h_P \tag{1.2} \\[6pt]
      \text{Covariance between two portfolios:}\quad & \operatorname{Cov}(r_P, r_Q) = \mathbf h_P^{T}\mathbf V\mathbf h_Q \tag{1.3} \\[6pt]
      \text{Exposure}\text{ of }P\text{ to }\mathbf a\text{:}\quad & a_P = \mathbf h_P^{T}\mathbf a = \mathbf a^{T}\mathbf h_P \tag{1.4}
    \end{align}

    <p>(1.2) and (1.3) are both linearity of covariance applied twice; see <a href="#bilinear">appendix&nbsp;C</a>.</p>

    <div class="def">
      <div class="label">Definition</div>
      <p style="margin:0">The <strong>characteristic portfolio of \(\mathbf a\)</strong>, written \(\mathbf h_a\), is the minimum-variance portfolio with unit exposure to \(\mathbf a\):</p>
      \[ \min_{\mathbf h}\ \sigma_{\mathbf h}^{2} \;=\; \min_{\mathbf h}\ \mathbf h^{T}\mathbf V\mathbf h \qquad \text{s.t.}\qquad \mathbf h^{T}\mathbf a = 1 \tag{1.5} \]
    </div>
    <p>No budget constraint \(\mathbf h^{T}\mathbf e = 1\), so \(\mathbf h_a\) need not be fully invested; and no long-only constraint \(h_n \ge 0\), so shorting and leverage are permitted.</p>

    <p><b>Subscripts on \(r\).</b> A subscript on \(r\) names a portfolio, so \(r_P = \mathbf h_P^{T}\mathbf r\) by (1.1); \(r_n\) with an asset index is the exception, meaning entry \(n\) of \(\mathbf r\). In particular \(r_a = \mathbf h_a^{T}\mathbf r\) is the return on the characteristic portfolio, not an asset return. Note that \(a_n\), asset \(n\)'s attribute value, is a different object from \(r_a\).</p>
  </section>

  <section id="foc">
    <h2><span class="num">02</span> First-Order Conditions via the Lagrangian</h2>

    <p>(1.5) is a constrained minimization problem. The Lagrangian converts it into a system of equations.</p>

    \[ L(\mathbf h,\theta) = \tfrac12\,\mathbf h^{T}\mathbf V\mathbf h - \theta\left(\mathbf h^{T}\mathbf a - 1\right) \tag{2.1} \]

    \[ \frac{\partial L}{\partial \theta} = 0 \;\Longrightarrow\; \mathbf a^{T}\mathbf h_a = 1 \tag{2.2} \]

    \begin{align}
      \frac{\partial L}{\partial \mathbf h} &= \mathbf 0 && \text{first-order condition in }\mathbf h \\[4pt]
      \frac{\partial}{\partial \mathbf h}\left[\tfrac12\,\mathbf h^{T}\mathbf V\mathbf h - \theta\left(\mathbf h^{T}\mathbf a - 1\right)\right] &= \mathbf 0 && \text{substituting (2.1)} \\[4pt]
      \tfrac12\left(2\,\mathbf V\mathbf h_a\right) - \theta\,\mathbf a &= \mathbf 0 && \text{by (A.2) and (A.1)} \\[4pt]
      \mathbf V\mathbf h_a &= \theta\,\mathbf a \tag{2.3}
    \end{align}

    <p>using \(\partial(\mathbf h^{T}\mathbf V\mathbf h)/\partial\mathbf h = 2\mathbf V\mathbf h\). See <a href="#calculus">appendix&nbsp;A</a> for that derivative. Because \(\mathbf V \succ 0\), the objective is strictly convex and the constraint is affine, so these conditions are not merely necessary but sufficient: they pin down the unique minimizer of (1.5), rather than a stationary point that might be a saddle.</p>

    <p>Note: \(\mathbf h_a\) has not been solved for. (2.3) leaves it inside a matrix product, and \(\theta\) is still unknown &mdash; so it fixes \(\mathbf h_a\)'s <em>direction</em>, parallel to \(\mathbf V^{-1}\mathbf a\), but not its length. The next two sections work with it: <a href="#reading">&#167;03</a> reads it as a condition on covariances and identifies \(\theta\) as the portfolio's own variance, and <a href="#beta">&#167;04</a> turns them into the beta identity. None of it inverts \(\mathbf V\). <a href="#altderiv">Appendix&nbsp;B</a> does the solving, two ways, for when the closed form is wanted.</p>
  </section>

  <section id="reading">
    <h2><span class="num">03</span> Covariance and Risk</h2>

<p>\(\mathbf V\mathbf h\) carries two interpretations: <b>marginal variance</b> (3.1) and <b>covariance with the portfolio</b> (3.2). Both hold for any portfolio, with no optimization involved. Imposing (2.3) then gives two results about the characteristic portfolio in particular: each asset's covariance with \(r_a\) is proportional to that asset's attribute (3.3), and \(\theta\) is the variance of \(r_a\) (3.4).</p>
<p><b>Marginal variance (3.1).</b> Differentiating (1.2) with respect to \(\mathbf h\) gives \(2\mathbf V\mathbf h\) (see <a href="#calculus">appendix&nbsp;A</a>). Entry \(n\) is</p>
    \[ \frac{\partial\sigma_P^{2}}{\partial h_{Pn}} = 2\left(\mathbf V\mathbf h_P\right)_n \tag{3.1} \]

    <p>the change in portfolio variance from holding a little more of asset \(n\). So \(\mathbf V\mathbf h_P\) is the vector of <strong>marginal risk contributions</strong>, up to the factor of 2.</p>

<p><b>Covariance with the portfolio (3.2).</b> Write out that same entry \((\mathbf V\mathbf h_P)_n\) as a sum, using linearity of covariance (<a href="#bilinear">appendix&nbsp;C</a>):</p>
    \[ \left(\mathbf V\mathbf h_P\right)_n = \sum_m V_{nm}h_{Pm} = \sum_m \operatorname{Cov}(r_n,r_m)\,h_{Pm} = \operatorname{Cov}\!\Big(r_n,\ \textstyle\sum_m h_{Pm} r_m\Big) = \operatorname{Cov}(r_n, r_P) \tag{3.2} \]

    <p>So \(\mathbf V\mathbf h_P\) is also the vector of <strong>covariances between each asset and the portfolio itself</strong>.</p>
    <p>Setting the two interpretations equal shows that an asset's marginal contribution to portfolio variance is proportional to its covariance with the portfolio:</p>

    \[ \frac{\partial\sigma_P^{2}}{\partial h_{Pn}} \;\propto\; \operatorname{Cov}(r_n, r_P) \]

    <p><b>Now impose (2.3).</b> Two independent things can be done with it. Component-wise it stays \(N\) equations, one per asset; contracted with \(\mathbf h_a\) it collapses to a single scalar. Neither derivation below uses the other.</p>

    <p><em>Entry by entry.</em></p>

    \begin{align}
      \operatorname{Cov}(r_n, r_a) &= \left(\mathbf V\mathbf h_a\right)_n && \text{(3.2) at }\mathbf h_P = \mathbf h_a \\[4pt]
      &= \left(\theta\,\mathbf a\right)_n && \text{by (2.3)} \\[4pt]
      &= \theta\,a_n && \text{for every asset }n \tag{3.3}
    \end{align}

    <p><em>Left-multiplied.</em> Multiply (2.3) through by \(\mathbf h_a^{T}\) and apply the constraint (2.2):</p>

    \begin{align}
      \mathbf V\mathbf h_a &= \theta\,\mathbf a && \text{(2.3)} \\[4pt]
      \mathbf h_a^{T}\mathbf V\mathbf h_a &= \theta\,\mathbf h_a^{T}\mathbf a && \text{left-multiply by }\mathbf h_a^{T} \\[4pt]
      \sigma_a^{2} &= \theta\cdot 1 && \text{by (1.2) and (2.2)} \\[4pt]
      \sigma_a^{2} &= \theta \tag{3.4}
    \end{align}

    <p>So the two conclusions are:</p>

    <ul class="plain">
      <li><b>(3.3)</b> &mdash; each asset's covariance with \(r_a\) is proportional to that asset's attribute, with the same constant \(\theta\) for every asset.</li>
      <li><b>(3.4)</b> &mdash; that constant \(\theta\) is the variance of \(r_a\).</li>
    </ul>
  </section>

  <section id="beta">
    <h2><span class="num">04</span> Exposure <em>is</em> beta</h2>

    <p>For any two returns beta is <em>defined</em> as the covariance ratio \(\beta_{X,Y} \equiv \operatorname{Cov}(X,Y)/\operatorname{Var}(Y)\) &mdash; the population slope of regressing \(X\) on \(Y\), equivalently the coefficient minimizing \(\mathbb E[(X-\beta Y)^{2}]\). Nothing about markets or equilibrium is assumed; it is a projection coefficient. Both arguments are returns, so a beta "against a portfolio" always means against that portfolio's return: \(\beta_{n,a}\) is shorthand for beta of \(r_n\) against \(r_a\), defined in <a href="#setup">&#167;01</a>.</p>

    <p><b>Asset level.</b> Divide (3.3) by \(\theta\). Then use (3.4), which says dividing by \(\theta\) is dividing by \(\sigma_a^{2}\), so the left side is exactly asset \(n\)'s beta:</p>

    \begin{align}
      \frac{\operatorname{Cov}(r_n, r_a)}{\theta} &= a_n && \text{(3.3), divided by }\theta \\[4pt]
      \frac{\operatorname{Cov}(r_n, r_a)}{\sigma_a^{2}} &= a_n && \text{by (3.4), }\theta = \sigma_a^{2} \\[4pt]
      \beta_{n,a} &= a_n \tag{4.1}
    \end{align}

    <p>Every asset's beta against \(r_a\) <em>is</em> its own attribute value.</p>

    <p><b>Portfolio level.</b> The same holds for any portfolio \(P\). Generalizing (4.1), apply the definition to \(r_P\) against \(r_a\), using the covariance form (1.3):</p>

    \[ \beta_{P,a} \equiv \frac{\operatorname{Cov}(r_P, r_a)}{\sigma_a^{2}} = \frac{\mathbf h_P^{T}\mathbf V\mathbf h_a}{\sigma_a^{2}} \tag{4.2} \]

    <p>(4.2) collapses two ways, asset by asset and in matrix form.</p>

    <p><em>Asset by asset.</em></p>

    \begin{align}
      \beta_{P,a} &= \frac{\operatorname{Cov}(r_P, r_a)}{\sigma_a^{2}} && \text{(4.2)} \\[4pt]
      &= \frac{1}{\sigma_a^{2}}\sum_n h_{Pn}\operatorname{Cov}(r_n, r_a) && \text{by (C.1)} \\[4pt]
      &= \sum_n h_{Pn}\,\frac{\operatorname{Cov}(r_n, r_a)}{\sigma_a^{2}} && \sigma_a^{2}\text{ is a constant} \\[4pt]
      &= \sum_n h_{Pn}\,\beta_{n,a} && \text{each term is an asset beta} \\[4pt]
      &= \sum_n h_{Pn}\,a_n && \text{by (4.1)} \\[4pt]
      &= a_P && \text{by (1.4)}
    \end{align}

    <p><em>In matrix form.</em> One step &mdash; do not expand \(\mathbf h_a\), use (2.3) directly:</p>

    \begin{align}
      \mathbf h_P^{T}\mathbf V\mathbf h_a &= \mathbf h_P^{T}\left(\theta\,\mathbf a\right) = \theta\,\mathbf h_P^{T}\mathbf a = \theta\,a_P && \text{by (2.3), then (1.4)} \\[4pt]
      \beta_{P,a} &= \frac{\theta\,a_P}{\sigma_a^{2}} = \frac{\theta\,a_P}{\theta} = a_P && \text{by (4.2), then (3.4)}
    \end{align}

    <div class="keybox">
      <div class="cap">Exposure = beta</div>
      \[ \beta_{P,a} = a_P \tag{4.3} \]
    </div>

    <p><b>Why this is not a coincidence.</b> In the geometry of <a href="#altderiv">appendix&nbsp;B</a>, \(\beta_{P,a}\) is the coefficient in the projection of \(\mathbf h_P\) onto \(\mathbf h_a\) measured in the covariance inner product, while \(a_P = \mathbf a^{T}\mathbf h_P\) is the plain product. (4.3) says these are the same number &mdash; which is identity (B.5), read again. The attribute \(\mathbf a\) measures exposure in the ordinary metric and \(\mathbf h_a\) measures it in the covariance metric; they agree because \(\mathbf h_a\) was built by converting the one into the other.</p>
  </section>

  <section id="ones">
    <h2><span class="num">05</span> The Global Minimum-Variance Portfolio</h2>

    <p>Take the attribute to be \(\mathbf e\), the vector of ones. Its exposure \(e_P = \mathbf h_P^{T}\mathbf e\) is total holdings, so <em>unit exposure means fully invested</em>, and the characteristic portfolio of \(\mathbf e\) — call it \(C\) — is the minimum-variance portfolio among all fully invested portfolios: the <strong>global minimum-variance portfolio</strong> &mdash; <em>global</em> meaning no expected-return target is imposed, only the budget constraint. Putting \(\mathbf a = \mathbf e\) into (B.2) and (B.3):</p>

    \[ \mathbf h_C = \frac{\mathbf V^{-1}\mathbf e}{\mathbf e^{T}\mathbf V^{-1}\mathbf e}, \qquad \sigma_C^{2} = \frac{1}{\mathbf e^{T}\mathbf V^{-1}\mathbf e} \tag{5.1} \]

    <p>And by (4.3), \(\beta_{P,C} = e_P = 1\) for <em>every</em> fully invested portfolio \(P\), whatever else it holds. Equivalently, straight from (2.3): \(\mathbf V\mathbf h_C = \sigma_C^{2}\,\mathbf e\) — \(C\) has the <em>same covariance with every asset</em>.</p>
  </section>

  <section id="why">
    <h2><span class="num">06</span> Why it matters</h2>

    <p>The map \(\mathbf a \mapsto \mathbf h_a\) defined by (1.5) is a dictionary, and \(\mathbf V\) does the translating: an attribute \(\mathbf a\) goes in, a holdings vector \(\mathbf h_a\) comes out, and \(\mathbf a\) reads back off that portfolio's covariances by (2.3). Attribute and holdings are the same object written two ways &mdash; a score per asset, or a position you can hold. Three reusable facts.</p>

    <ol class="plain">
      <li><strong>\(\mathbf h_a \propto \mathbf V^{-1}\mathbf a\).</strong> Every minimum-risk-per-unit-of-something portfolio has this shape; only the normalizer changes.</li>
      <li><strong>The Lagrange multiplier <em>is</em> the variance.</strong> \(\theta = \sigma_a^{2}\) by (3.4), and (B.3) evaluates it, so \(1/(\mathbf a^{T}\mathbf V^{-1}\mathbf a)\) is a risk number rather than just algebra.</li>
      <li><strong>Exposure to an attribute equals beta against that attribute's characteristic portfolio.</strong> </li>
    </ol>

    <p>The third fact means any cross-sectional characteristic can be restated as a beta. Anything you can say about betas (hedging, attribution, decomposition) you can say about characteristics. Say the attribute is book-to-price and the mandate is no net value tilt, \(a_P = 0\). By (4.3) that is the same requirement as \(\beta_{P,a} = 0\): neutrality to a score becomes an ordinary hedge against one portfolio, \(\mathbf h_a\), which can be sized and traded like any other.</p>
  </section>

  <section id="assumptions">
    <h2><span class="num">07</span> Assumptions</h2>

    <ol class="plain caveats">
      <li><b>\(\mathbf V\) positive definite.</b> A covariance matrix is only guaranteed PSD; PD is what makes \(\mathbf V^{-1}\) exist and the minimizer unique. It fails outright once \(N\) exceeds the sample length.</li>
      <li><b>\(\mathbf a \neq \mathbf 0\)</b> for feasibility, and \(\mathbf a^{T}\mathbf V^{-1}\mathbf a &gt; 0\) for (B.1) — the second follows from the first plus PD.</li>
      <li><b>The sign convention on \(\theta\).</b> Writing the constraint term as \(-\theta(\mathbf h^{T}\mathbf a - 1)\) is what makes \(\theta\) come out positive and equal to \(\sigma_a^{2}\); the opposite sign flips it.</li>
      <li><b>No budget or sign constraints.</b> \(\mathbf h_a\) need not sum to 1 and need not be long-only.</li>
      <li><b>Beta here is pure covariance algebra.</b> No equilibrium, no expected returns, no CAPM content. </li>
    </ol>
  </section>


  <section id="calculus">
    <h2><span class="num">A</span> Appendix — the matrix derivatives used</h2>
    <p>Two derivatives are all the matrix calculus (2.3) needs. Each is a one-liner once written entry by entry, which is also the quickest way to re-derive them cold rather than recall them.</p>

    <p><b>Layout convention.</b> For scalar \(f\) and column vector \(\mathbf x\) (\(N\times1\)), \(\partial f/\partial\mathbf x\) is the \(N\times1\) column of partials \((\partial f/\partial x_k)\).</p>

    <p><b>(A.1) Linear form.</b> \(f = \mathbf b^{T}\mathbf x = \sum_i b_i x_i\), so \(\partial f/\partial x_k = b_k\). A scalar equals its own transpose, so \(\mathbf x^{T}\mathbf b = \mathbf b^{T}\mathbf x\) and the two orderings differentiate alike:</p>

    \[ \frac{\partial}{\partial\mathbf x}\left(\mathbf b^{T}\mathbf x\right) = \frac{\partial}{\partial\mathbf x}\left(\mathbf x^{T}\mathbf b\right) = \mathbf b \tag{A.1} \]

    <p><b>(A.2) Quadratic form.</b> \(f = \mathbf x^{T}\mathbf A\mathbf x = \sum_i\sum_j x_i A_{ij} x_j\). Differentiating in \(x_k\) hits the sum twice — once through \(i=k\), once through \(j=k\):</p>

    \[ \frac{\partial f}{\partial x_k} = \sum_j A_{kj}x_j + \sum_i x_i A_{ik} = \left(\mathbf A\mathbf x\right)_k + \left(\mathbf A^{T}\mathbf x\right)_k \]

    \[ \frac{\partial}{\partial\mathbf x}\left(\mathbf x^{T}\mathbf A\mathbf x\right) = \left(\mathbf A + \mathbf A^{T}\right)\mathbf x \;=\; 2\mathbf A\mathbf x \quad\text{when } \mathbf A = \mathbf A^{T} \tag{A.2} \]

    <p>For general \(\mathbf A\) the factor is \((\mathbf A+\mathbf A^{T})\), and only for symmetric \(\mathbf A\) does it collapse to \(2\mathbf A\mathbf x\). This is the step where "\(\mathbf V\) symmetric" is spent.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:24px 0 8px">Applying them to (2.3)</h3>

    \[ L(\mathbf h,\theta) = \underbrace{\tfrac12\,\mathbf h^{T}\mathbf V\mathbf h}_{\text{(A.2)},\ \mathbf A=\mathbf V} - \theta\big(\underbrace{\mathbf h^{T}\mathbf a}_{\text{(A.1)},\ \mathbf b=\mathbf a} - 1\big) \]

    \[ \frac{\partial L}{\partial\mathbf h} = \tfrac12\left(2\mathbf V\mathbf h\right) - \theta\,\mathbf a = \mathbf V\mathbf h - \theta\,\mathbf a \]

    \[ \frac{\partial L}{\partial\mathbf h} = \mathbf 0 \;\Longrightarrow\; \mathbf V\mathbf h_a = \theta\,\mathbf a \tag{2.3} \]

    <p>The \(-1\) is constant in \(\mathbf h\) and drops; the \(\tfrac12\) cancels the 2 from (A.2), which is all that convention is for.</p>


    <p class="note" style="border-top:1px dashed var(--rule)"><b>Two transpose facts used elsewhere.</b> \(\mathbf V\) symmetric \(\Rightarrow\) \(\mathbf V^{-1}\) symmetric: transpose \(\mathbf V\mathbf V^{-1} = \mathbf I\) to get \((\mathbf V^{-1})^{T}\mathbf V^{T} = (\mathbf V^{-1})^{T}\mathbf V = \mathbf I\), so \((\mathbf V^{-1})^{T} = \mathbf V^{-1}\). And \((\mathbf A\mathbf B)^{T} = \mathbf B^{T}\mathbf A^{T}\) gives \((\mathbf V^{-1}\mathbf a)^{T} = \mathbf a^{T}\mathbf V^{-1}\). </p>

  </section>

  <section id="altderiv">
    <h2><span class="num">B</span> Appendix &mdash; solving for \(\mathbf h_a\), two ways</h2>

    <p>Nothing in the main line needs an explicit \(\mathbf h_a\): &#167;03 and &#167;04 run on (2.3) alone. The closed form is still worth having, and deriving it twice separates what is calculus from what is geometry.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:26px 0 10px">Method 1 &mdash; elimination, from the first-order conditions</h3>

    <p>&#167;02 produced two conditions and nothing else: (2.3), which is \(N\) equations, and (2.2), which is one. The unknowns are the \(N\) components of \(\mathbf h_a\) plus the scalar \(\theta\) &mdash; a square system, \(N+1\) by \(N+1\), and the rest is elimination.</p>

    <p><b>Step 1 &mdash; eliminate \(\mathbf h_a\).</b> \(\mathbf V\) is invertible, so (2.3) gives \(\mathbf h_a = \theta\,\mathbf V^{-1}\mathbf a\). Putting that into (2.2) leaves \(\theta\) as the only unknown:</p>

    \[ \theta\,\mathbf a^{T}\mathbf V^{-1}\mathbf a = 1 \qquad\Longrightarrow\qquad \theta = \frac{1}{\mathbf a^{T}\mathbf V^{-1}\mathbf a} \tag{B.1} \]

    <p><b>Step 2 &mdash; substitute \(\theta\) back.</b> Putting (B.1) into \(\mathbf h_a = \theta\,\mathbf V^{-1}\mathbf a\) removes \(\theta\):</p>

    <div class="keybox">
      <div class="cap">Characteristic portfolio of \(\mathbf a\)</div>
      \[ \mathbf h_a = \frac{\mathbf V^{-1}\mathbf a}{\mathbf a^{T}\mathbf V^{-1}\mathbf a} \;=\; \theta\,\mathbf V^{-1}\mathbf a \tag{B.2} \]
    </div>

    <p>The <em>direction</em> of \(\mathbf h_a\) is \(\mathbf V^{-1}\mathbf a\); the constraint only fixes its length, and the denominator is exactly the normalizer that sets \(\mathbf a^{T}\mathbf h_a = 1\).</p>

    <p><b>Step 3 &mdash; read off the risk.</b> (3.4) already gave \(\sigma_a^{2} = \theta\), and (B.1) gives \(\theta\):</p>

    <div class="keybox">
      <div class="cap">Variance of \(\mathbf h_a\)</div>
      \[ \sigma_a^2 = \theta = \frac{1}{\mathbf a^{T}\mathbf V^{-1}\mathbf a}, \qquad \sigma_a = \left(\mathbf a^{T}\mathbf V^{-1}\mathbf a\right)^{-1/2} \tag{B.3} \]
    </div>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:30px 0 10px">Method 2 &mdash; without calculus</h3>

    <p>Method 1 <em>produces</em> \(\mathbf V^{-1}\mathbf a\) but does not explain it. This route gets the same answer from linear algebra alone, and says what \(\mathbf V^{-1}\mathbf a\) is: the constraint plane's normal direction, measured in the geometry \(\mathbf V\) defines.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:24px 0 8px">Step 1 &mdash; \(\mathbf V\) defines a geometry in which length is risk</h3>

    <p>Because \(\mathbf V \succ 0\), the form \(\langle\mathbf x,\mathbf y\rangle_V \equiv \mathbf x^{T}\mathbf V\mathbf y\) satisfies the three axioms of an inner product: it is symmetric because \(\mathbf V = \mathbf V^{T}\), bilinear because matrix multiplication is, and positive definite because \(\mathbf x^{T}\mathbf V\mathbf x > 0\) for \(\mathbf x \neq \mathbf 0\). Every inner product induces a length by \(\lVert\mathbf x\rVert^{2} = \langle\mathbf x,\mathbf x\rangle\). Here that length is exactly the portfolio's risk:</p>

    \[ \lVert\mathbf h\rVert_V^{2} = \langle\mathbf h,\mathbf h\rangle_V = \mathbf h^{T}\mathbf V\mathbf h = \sigma_{\mathbf h}^{2} \tag{B.4} \]

    <p>So the problem (1.5) reads: among the feasible portfolios, find the shortest one.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:24px 0 8px">Step 2 &mdash; rewrite the constraint in that same geometry</h3>

    <p>The objective is now stated with \(\langle\cdot,\cdot\rangle_V\), but the constraint \(\mathbf h^{T}\mathbf a = 1\) is still stated with the ordinary product. Put them in the same terms by inserting \(\mathbf V\mathbf V^{-1} = \mathbf I\):</p>

    \[ \mathbf h^{T}\mathbf a = \mathbf h^{T}\mathbf V\mathbf V^{-1}\mathbf a = \big\langle\, \mathbf h,\ \mathbf V^{-1}\mathbf a \,\big\rangle_V \tag{B.5} \]

    <p><b>This is the step that produces \(\mathbf V^{-1}\mathbf a\)</b>, and it is worth saying what it means. A plane \(\{\mathbf h : \mathbf h^{T}\mathbf a = 1\}\) has a normal direction, but "normal" depends on which inner product you measure with. Under the ordinary product the normal is \(\mathbf a\); under \(\langle\cdot,\cdot\rangle_V\) it is \(\mathbf n \equiv \mathbf V^{-1}\mathbf a\). The plane is the same set of portfolios either way &mdash; only the description changes, and the objective measures with \(\langle\cdot,\cdot\rangle_V\), so \(\mathbf n\) is the description that matters.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:24px 0 8px">Step 3 &mdash; the problem is now standard</h3>

    <p>Substituting (B.4) and (B.5), the problem is</p>

    \[ \min_{\mathbf h}\ \lVert\mathbf h\rVert_V \qquad\text{s.t.}\qquad \langle\mathbf h,\mathbf n\rangle_V = 1 \]

    <p>Minimizing \(\lVert\mathbf h\rVert_V\) rather than \(\lVert\mathbf h\rVert_V^{2}\) changes nothing, since \(t \mapsto t^{2}\) is increasing on \(t \ge 0\) and so has the same minimizer.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:24px 0 8px">Step 4 &mdash; Cauchy&ndash;Schwarz bounds the objective from below</h3>

    <p>In any inner product space, \(|\langle\mathbf x,\mathbf y\rangle| \le \lVert\mathbf x\rVert\,\lVert\mathbf y\rVert\), with equality if and only if the two vectors are parallel. Apply it to any feasible \(\mathbf h\), whose inner product against \(\mathbf n\) is fixed at 1:</p>

    \[ 1 = \langle\mathbf h,\mathbf n\rangle_V \;\le\; \lVert\mathbf h\rVert_V\,\lVert\mathbf n\rVert_V \qquad\Longrightarrow\qquad \lVert\mathbf h\rVert_V \;\ge\; \frac{1}{\lVert\mathbf n\rVert_V} \]

    <p>So \(1/\lVert\mathbf n\rVert_V\) is a floor on risk that no feasible portfolio can beat. Nothing yet says it is reached.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:24px 0 8px">Step 5 &mdash; the floor is attained, by exactly one portfolio</h3>

    <p>Step 4 gives a floor, but nothing yet says any portfolio reaches it. Suppose some feasible \(\mathbf h\) sits exactly at it, \(\lVert\mathbf h\rVert_V = 1/\lVert\mathbf n\rVert_V\). Then</p>

    \begin{align}
      \lVert\mathbf h\rVert_V\,\lVert\mathbf n\rVert_V &= \frac{1}{\lVert\mathbf n\rVert_V}\cdot\lVert\mathbf n\rVert_V = 1 && \text{it sits at the floor} \\[4pt]
      \langle\mathbf h,\mathbf n\rangle_V &= 1 && \text{it is feasible} \\[4pt]
      \langle\mathbf h,\mathbf n\rangle_V &= \lVert\mathbf h\rVert_V\,\lVert\mathbf n\rVert_V && \text{so the two agree}
    \end{align}

    <p>Cauchy&ndash;Schwarz has become an equality, which happens only for parallel vectors. So any \(\mathbf h\) at the floor must be a multiple \(\mathbf h = c\,\mathbf n\), and the constraint then fixes \(c\):</p>

    \begin{align}
      \langle c\,\mathbf n,\ \mathbf n\rangle_V &= 1 && \text{the constraint} \\[4pt]
      c\,\lVert\mathbf n\rVert_V^{2} &= 1 && \text{scalars come out of the inner product} \\[4pt]
      c &= \frac{1}{\lVert\mathbf n\rVert_V^{2}}
    \end{align}

    <p>So there is exactly one candidate, \(\mathbf h_a = \mathbf n/\lVert\mathbf n\rVert_V^{2}\). Check that it does reach the floor:</p>

    \[ \lVert\mathbf h_a\rVert_V = \frac{\lVert\mathbf n\rVert_V}{\lVert\mathbf n\rVert_V^{2}} = \frac{1}{\lVert\mathbf n\rVert_V} \]

    <p>It is feasible and it attains the lower bound, so it is a minimizer. And it is the only one: reaching the floor required being parallel to \(\mathbf n\), and the constraint left just one multiple.</p>

    <h3 style="font-family:'Space Grotesk',sans-serif;font-weight:600;font-size:15px;margin:24px 0 8px">Step 6 &mdash; evaluate</h3>

    <p>Only \(\lVert\mathbf n\rVert_V^{2}\) is left to compute, and the two \(\mathbf V\)'s cancel against the \(\mathbf V^{-1}\)'s:</p>

    \[ \lVert\mathbf n\rVert_V^{2} = \big\langle \mathbf V^{-1}\mathbf a,\ \mathbf V^{-1}\mathbf a\big\rangle_V = \mathbf a^{T}\mathbf V^{-1}\mathbf V\mathbf V^{-1}\mathbf a = \mathbf a^{T}\mathbf V^{-1}\mathbf a \]

    <p>Substituting back reproduces both earlier results, with no Lagrangian and no derivatives anywhere:</p>

    \[ \mathbf h_a = \frac{\mathbf V^{-1}\mathbf a}{\mathbf a^{T}\mathbf V^{-1}\mathbf a} \;\;\text{(B.2)}, \qquad \sigma_a = \lVert\mathbf h_a\rVert_V = \frac{1}{\lVert\mathbf n\rVert_V} = \big(\mathbf a^{T}\mathbf V^{-1}\mathbf a\big)^{-1/2} \;\;\text{(B.3)} \]

    </section>

  <section id="bilinear">
    <h2><span class="num">C</span> Appendix &mdash; covariance is linear in holdings</h2>

    <p>For deterministic weights \(c_n\) and any return \(Y\), start from the definition of covariance and use linearity of expectation twice:</p>

    \begin{align}
      \operatorname{Cov}\!\Big(\sum_n c_n r_n,\ Y\Big)
        &= \mathbb E\Big[\Big(\sum_n c_n r_n - \mathbb E\Big[\sum_n c_n r_n\Big]\Big)\left(Y - \mathbb EY\right)\Big] && \text{definition} \\[4pt]
        &= \mathbb E\Big[\Big(\sum_n c_n\left(r_n - \mathbb Er_n\right)\Big)\left(Y - \mathbb EY\right)\Big] && \text{center; }c_n\text{ constant} \\[4pt]
        &= \mathbb E\Big[\sum_n c_n\left(r_n - \mathbb Er_n\right)\left(Y - \mathbb EY\right)\Big] && \text{multiply through} \\[4pt]
        &= \sum_n c_n\,\mathbb E\big[\left(r_n - \mathbb Er_n\right)\left(Y - \mathbb EY\right)\big] && \text{linearity of }\mathbb E \\[4pt]
        &= \sum_n c_n \operatorname{Cov}(r_n, Y) && \text{definition} \tag{C.1}
    \end{align}

    <p>The same holds in the second argument by the symmetry of covariance, and applying it in both gives (1.2) and (1.3). Only linearity of expectation is used &mdash; no independence and no distributional assumption &mdash; but the weights must be deterministic. If holdings responded to returns, \(c_n\) could not leave the expectation.</p>
  </section>

</div>

<nav class="pagenav foot"><a href="/">&larr; Mike Purewal</a> &nbsp;&middot;&nbsp; <a href="/writing/">All writing</a> &nbsp;&middot;&nbsp; <a href="/tag/grinold-kahn">More in this series</a></nav>

</body>
</html>]]></content><author><name></name></author><category term="grinold-kahn" /><category term="finance" /><category term="mathematics" /><summary type="html"><![CDATA[Characteristic Portfolios]]></summary></entry><entry><title type="html">The NBA Salary Cap</title><link href="/2026/08/29/nba-salary-cap.html" rel="alternate" type="text/html" title="The NBA Salary Cap" /><published>2026-08-29T00:00:00+00:00</published><updated>2026-08-29T00:00:00+00:00</updated><id>/2026/08/29/nba-salary-cap</id><content type="html" xml:base="/2026/08/29/nba-salary-cap.html"><![CDATA[<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>The NBA Salary Cap</title>
<!--
  Source: github.com/meninder/learn-nba-cap-rules -- synthesis/one-page.html
  Copied 2026-08-29 from commit dc7a329.
  GENERATED FILE -- do not edit here. The source repo is the source of truth.
  Edit synthesis/one-page.html there, then regenerate this file's body and
  assets/css/nba-cap.css. Both are byte-identical to the source, so diff is
  the check that they are in sync.
-->
<link rel="icon" type="image/png" href="/favicon.png">
<link rel="stylesheet" href="/assets/css/nba-cap.css"><script async src="https://www.googletagmanager.com/gtag/js?id=G-9X2510WP61"></script>
<script>
  window.dataLayer = window.dataLayer || [];
  function gtag(){dataLayer.push(arguments);}
  gtag('js', new Date());

  gtag('config', 'G-9X2510WP61');
</script></head>
<body>

<h1>The NBA Salary Cap</h1>

<p class="standfirst">A framework for following the news, because player movement is run by a rulebook. Every move a team makes is one of three things: retaining its own player, signing a new one, or trading. Where the team sits on the ladder below decides which of the three are still open to it.</p>

<h2>Introduction</h2>

<p>Salary cap rules live in the <strong>Collective Bargaining Agreement</strong>, the contract between the league and the National Basketball Players Association. The current one was signed in 2023. It sets the thresholds, the exceptions, and the penalties, and it is renegotiated periodically.</p>

<p>The cap itself moves annually because it is tied to league revenue. Figures below are 2025-26. </p>

<div class="table-wrap">
<table class="data">
  <tr><th>Rung</th><th>2025-26</th><th>What it does</th></tr>
  <tr><td>Salary floor</td><td class="fig">$139.2M</td><td>Minimum spend. Shortfalls go to your players anyway.</td></tr>
  <tr><td>Salary cap</td><td class="fig">$154.6M</td><td>Crossable only through named exceptions.</td></tr>
  <tr><td>Luxury tax</td><td class="fig">$187.9M</td><td>Every dollar over is taxed, in escalating brackets.</td></tr>
  <tr><td>First apron</td><td class="fig">$195.9M</td><td>Roster-building tools start being removed.</td></tr>
  <tr><td>Second apron</td><td class="fig">$207.8M</td><td>Most remaining tools gone; draft picks attacked.</td></tr>
</table>
</div>

<p>It is a <strong>soft cap</strong>. You may cross the cap line, but only through specific authorizations called <strong>exceptions</strong>. There is no general permission to exceed it, only a list.</p>

<div class="key">
  <span class="label">One total, two states</span>
  <p>Every contract on the books counts toward <strong>one</strong> team salary total, and that number puts a team in one of two states.</p>
  <p><strong>Under the cap</strong>: real space to sign anyone, but the big exceptions are forfeited.</p>
  <p><strong>Over the cap</strong>: no space at all, and players are added only through exceptions or trades.</p>
</div>

<h2>1. Over the cap</h2>

<p>The most common condition. A team does three things from here: keep its own players, sign new ones, and trade for them. Keeping is nearly unrestricted. Signing and trading are governed by two separate systems. </p>

<div class="key">
  <span class="label">Two engines</span>
  <p><strong>Signing</strong> runs on exceptions: named permissions to add salary while over the cap.</p>
  <p><strong>Trading</strong> runs on salary matching: what you send out must roughly cover what you take back.</p>
</div>

<h3>1A. Retaining your own players</h3>

<p>A hard cap would force teams to release their own stars the moment they got expensive. <strong>Bird rights</strong><sup class="fn"><a href="#fn1" id="fnref1">1</a></sup> exist to prevent that: they let a team exceed the cap, the tax, <em>and both aprons</em> for one purpose only, re-signing a player already on its roster. Retention is the one thing the entire penalty system never touches.</p>

<p>They come in three tiers scaled to how long the player has been yours.</p>

<div class="table-wrap">
<table class="data">
  <tr><th>Tier</th><th>Earned after</th><th>Can pay up to</th><th>Length</th></tr>
  <tr><td>Full Bird</td><td>3 seasons</td><td>The max salary, 8% raises</td><td>5 years</td></tr>
  <tr><td>Early Bird</td><td>2 seasons</td><td>Greater of 175% of last salary or the league average</td><td>2 to 4 years</td></tr>
  <tr><td>Non-Bird</td><td>1 season</td><td>Greater of 120% of last salary or 120% of the minimum</td><td>4 years</td></tr>
</table>
</div>

<p>So retention is unrestricted for a player you have had three years, and constrained below that.</p>

<p>That constraint has teeth. Take an unrestricted free agent who arrived two seasons ago: his own team holds only Early Bird rights, so the most it can offer him is near the league average, roughly mid-level money, however much he is now worth. The permission to exceed the cap is real, but below the third season it stops well short of the max. The escape hatch is restricted free agency, where the right to match an offer sheet works at any payroll and needs no exception at all.</p>

<p><strong>The clock asymmetry.</strong> Bird rights travel with a player in a trade, so the acquiring team inherits his accumulated seasons. They reset to zero if he changes teams as a free agent. This asymmetry is the entire reason sign-and-trades exist.</p>

<p><strong>Restricted free agency.</strong> A second retention tool. Extend a young player a <em>qualifying offer</em> and you hold the right to match any offer sheet he signs elsewhere. Heavy team leverage, and the reason these standoffs run for months.</p>

<h3>1B. Signing new players</h3>

<p>Three tools, in descending order of size.</p>

<div class="table-wrap">
<table class="data">
  <tr><th>Tool</th><th>Max starting salary</th><th>Length</th><th>Notes</th></tr>
  <tr><td>Mid-level exception (MLE)</td><td class="fig">$14.1M</td><td>4 yrs</td><td>One mid-sized slot per year. The primary way a capped-out contender adds a real rotation player.</td></tr>
  <tr><td>Bi-annual exception (BAE)</td><td class="fig">~$5M</td><td>2 yrs</td><td>Usable only in alternating years.</td></tr>
  <tr><td>Minimum contract</td><td class="fig">League min</td><td>up to 2 yrs</td><td>Always available to every team at any payroll. No team is ever unable to sign anyone.</td></tr>
</table>
</div>

<p>The figures are first-year salaries; raises are added on top over the life of the deal. These are the full-strength versions, held by a team over the cap but below the first apron. Each apron shrinks or removes them, which is the subject of the penalties section.</p>

<h3>1C. Trading</h3>

<p>Trades use <strong>salary matching</strong>: what you send out has to roughly cover what you take back.</p>

<h4>Matching</h4>

<ul class="rules">
  <li>The band is <strong>125% of outgoing salary plus $100K</strong>. Smaller salaries get more generous multipliers, up to 200%, but 125% governs the trades that make news. This is why marginal players get attached to deals, purely to make the arithmetic legal.</li>
  <li>It is an <strong>over-the-cap rule only</strong>. Two under-cap teams can swap wildly mismatched salaries.</li>
  <li>It caps what you take <strong>in</strong>, never what you send <strong>out</strong>. Shedding salary is always legal and never needs matching. In a lopsided deal the burden falls entirely on the side receiving the larger number.</li>
</ul>

<h4>Aggregation</h4>

<p>Combining two or more outgoing contracts so their sum matches one larger incoming contract: three role players out, one star back. No single contract on the roster is big enough to cover a $40M star, and the band demands roughly $32M going the other way, so the team bundles three deals of about $11M each until the sum clears it. This is the everyday machinery of blockbuster trades.</p>

<h4>Traded player exception (TPE)</h4>

<p>Send out more salary than you take back and the difference is banked as a credit. Shed $30.7M, take back $8.2M, bank roughly $22.5M. The TPE absorbs one incoming contract in a <em>later</em> trade with nothing sent out to match it. It is homemade cap space for a team that has none.</p>

<ul class="rules">
  <li><strong>Size:</strong> salary shed, plus a cushion of about $250K if you are below the first apron. No cushion above it.</li>
  <li><strong>Expiry:</strong> one year. Use it or lose it.</li>
  <li><strong>One at a time:</strong> it absorbs a single salary that fits underneath it. You cannot combine two TPEs, and you cannot stack other salary on top of one.</li>
</ul>

<h4>Sign-and-trade</h4>

<p>A star wants to leave Team A for Team B, but B is over the cap and can only offer him the mid-level. Signing outright costs him tens of millions and leaves A with nothing. So A re-signs its own free agent first, using Bird rights to pay him well past what B could offer, then trades that contract to B under normal matching. The Bird-sized value rides along with the player.</p>

<div class="table-wrap">
<table class="data">
  <tr><th>Party</th><th>What it gets</th></tr>
  <tr><td>The player</td><td>A Bird-sized contract instead of a mid-level one, on the team he wanted.</td></tr>
  <tr><td>The old team</td><td>Players and picks instead of nothing.</td></tr>
  <tr><td>The new team</td><td>A player it had no room to sign.</td></tr>
</table>
</div>

<p>Nothing in it is a new rule. It is Bird rights and salary matching arranged in sequence. Matching still applies to the trade leg, and the resulting contract must run either 3 or 4 years, with only the first year fully guaranteed.</p>

<p>That length cap is deliberate. Before 2011 a sign-and-trade could run five years, which made leaving as lucrative as staying. The current rule holds sign-and-trade deals to the same four-year ceiling a team with cap room faces, so that the five-year contract remains available only from your own team, re-signing outright.<sup class="fn"><a href="#fn2" id="fnref2">2</a></sup> The financial incentive to stay put is the point.</p>

<h2>2. Under the cap</h2>

<p>Rare. Entering 2025 free agency, exactly one team, Brooklyn, had meaningful cap space, roughly $35M. A couple of others could manufacture some by giving things up. The rest of the league was over the cap. Seven teams used room in 2024.</p>

<h3>2A. What you get</h3>

<ul class="rules">
  <li>Sign any free agent outright, up to the amount of room you have. No exception needed.</li>
  <li>Absorb salary in trades without matching. You just need room to fit what is arriving.</li>
</ul>

<h3>2B. What it costs</h3>

<ul class="rules">
  <li><strong>Room runs out.</strong> Space is not a standing privilege, it is a balance you draw down. Sign to the limit of it and you are an over-the-cap team for the rest of the summer, back to needing an exception to add anyone.</li>
  <li><strong>And the big exception is forfeited.</strong> A team that used room does not hold the $14.1M mid-level. What it has for the rest of the offseason is the smaller <strong>room exception</strong>: $8.8M over 3 years. You cannot hold both at full value, so the choice is one large signing now against a stronger follow-up later.</li>
  <li><strong>Cap holds.</strong> Your own outgoing free agents occupy a placeholder charge against team salary until you re-sign them or renounce them. To create real room you must renounce, and renouncing destroys the Bird rights that would have let you exceed the cap to keep that player. Space is bought with retention.</li>
</ul>

<p>Spend the room and you are an over-the-cap team with a worse toolbox.</p>

<h2>3. Penalties</h2>

<p>Everything above describes a team's full toolbox. This section is what gets taken back as payroll climbs.</p>

<div class="key">
  <span class="label">Three currencies</span>
  <p>Low on the ladder you pay nothing. In the middle, above the luxury tax line, you pay <strong>money</strong>. At the <strong>first apron</strong> you stop paying money and start paying in <strong>tools</strong>, and at the <strong>second apron</strong> in <strong>draft picks</strong>. The change of currency is the whole design. A fine you can afford is a price, and the owners the league most wanted to constrain were exactly the ones who could pay it.</p>
</div>

<h3>3A. Money: the luxury tax</h3>

<p>Above $187.9M every dollar is taxed, in brackets that grow steeper as you climb. A team that has paid the tax in three of the previous four seasons pays a higher <strong>repeater</strong> multiplier at every bracket. The compounding is severe enough that shedding modest salary can save several times that amount in tax.</p>

<h3>3B. Tools: the aprons</h3>

<div class="key">
  <span class="label">The organizing rule</span>
  <p>Every apron restriction blocks bringing in <em>outside</em> talent. Bird rights, the tool for keeping your own, is untouched at every level. The message is not "you cannot be expensive." It is: stay as expensive as you like keeping the team you built, but you may no longer shop for more.</p>
</div>

<p>Here is the toolbox shrinking. This single table is the apron system.</p>

<div class="table-wrap">
<table class="data">
  <tr>
    <th>Tool</th>
    <th>Over cap, below apron 1</th>
    <th>Over first apron</th>
    <th>Over second apron</th>
  </tr>
  <tr><td>Bird rights</td><td>Full</td><td>Full</td><td>Full</td></tr>
  <tr><td>Mid-level exception</td><td class="fig">$14.1M / 4 yrs</td><td class="fig">$5.7M / 2 yrs</td><td class="no">None</td></tr>
  <tr><td>Bi-annual exception</td><td class="fig">~$5M</td><td class="no">Gone</td><td class="no">Gone</td></tr>
  <tr><td>Minimum contracts</td><td>Yes</td><td>Yes</td><td>Yes</td></tr>
  <tr><td>Salary matching</td><td class="fig">125% + $100K</td><td class="fig">110%</td><td class="no">No more than sent</td></tr>
  <tr><td>Aggregation</td><td>Yes</td><td>Yes</td><td class="no">Banned</td></tr>
  <tr><td>Prior-year TPEs</td><td>Usable</td><td class="no">Unusable</td><td class="no">Unusable</td></tr>
  <tr><td>Sign-and-trade in</td><td>Yes</td><td class="no">Banned</td><td class="no">Banned</td></tr>
  <tr><td>Buyout signings</td><td>Any player</td><td class="no">Not if his old salary topped ~$12.4M</td><td class="no">Not if his old salary topped ~$12.4M</td></tr>
  <tr><td>Cash in trades</td><td>Yes</td><td>Yes</td><td class="no">Banned</td></tr>
</table>
</div>

<p>Read down the MLE row: a $14.1M offer to an outside free agent collapses to $5.7M, then to nothing. Above the second apron your entire pitch to an outsider is a minimum contract, however much the owner is willing to spend.</p>

<p>Read the second-apron column and the difference is one of kind, not degree. The first apron takes away ways to <em>sign</em> people. The second apron strips the machinery of <em>trading itself</em>. With aggregation banned you would need a single outgoing player large enough to match a star on his own, and almost no team has one. Functionally, a second-apron team cannot trade for a star.</p>

<h3>3C. Picks: the future</h3>

<ul class="rules">
  <li><strong>Frozen pick.</strong> A team above the second apron cannot trade its first-round pick seven years out.</li>
  <li><strong>The three-of-five penalty.</strong> Finish above the second apron in three of any five seasons and that pick is moved to <strong>No. 30</strong>, last in the first round, regardless of record. A franchise that spends heavily, ages, and collapses can earn the third pick on the floor and be required to draft thirtieth.</li>
</ul>

<p>Together these do not force a breakup. They make staying together self-defeating, by removing the two things a team needs to stay good: trade flexibility and draft capital. Hence the nickname "dynasty killer."</p>

<h3>3D. Getting back under</h3>

<ul class="rules">
  <li><strong>Trade salary away.</strong> Always legal, never needs matching.</li>
  <li><strong>The stretch provision.</strong> Waive a player and spread his remaining guaranteed money over (2 x years left) + 1 seasons, shrinking the annual cap hit. You get under a line at the cost of paying someone not to play for years.</li>
  <li><strong>Buyouts.</strong> Team and player agree to terminate the contract, usually with the player returning some money.</li>
</ul>

<h3>3E. The mechanism in action</h3>

<p>In the summer of 2025 Boston traded Jrue Holiday and Kristaps Porzingis, two starters from a team that had won the title fourteen months earlier. Jayson Tatum had ruptured an Achilles and the season was likely lost. Ownership could pay the bill, reported near $500 million with tax included. What they would not absorb was what came attached to it. Brad Stevens, their president of basketball operations:</p>

<blockquote>"The second apron is why those trades happen. The basketball penalties associated with those are real."</blockquote>

<p>Both were straight salary-matched trades. No exception was involved anywhere, because trades do not run on exceptions. Moving Porzingis at roughly $30.7M for Georges Niang at roughly $8.2M is what dropped them under the second apron, and because they were deep in repeater territory the reported tax saving was around $40M on a salary cut of about $22.5M.</p>

<h2>4. Other</h2>

<h3>4A. Hard caps</h3>

<p>The NBA is a soft cap, but certain moves bolt a temporary, self-inflicted <strong>hard cap</strong> onto the team that makes them: a line it cannot cross for the rest of the season, no matter what Bird rights it holds.</p>

<div class="table-wrap">
<table class="data">
  <tr><th>Do this</th><th>Hard cap lands at</th></tr>
  <tr><td>Use the full non-taxpayer MLE</td><td>First apron</td></tr>
  <tr><td>Use the bi-annual exception</td><td>First apron</td></tr>
  <tr><td>Acquire a player via sign-and-trade</td><td>First apron</td></tr>
  <tr><td>Use the taxpayer MLE</td><td>Second apron</td></tr>
</table>
</div>

<p>So there are two different ways to be at an apron: climb over it and live under the restrictions, or trigger a wall from below by reaching for a big tool. The logic is consistent. The tools the first apron forbids from above are exactly the ones that pin you underneath it from below.</p>

<h3>4B. Compression</h3>

<p>The whole system collapses to one question: <strong>how do I legally absorb incoming salary?</strong></p>

<ul class="rules">
  <li><strong>Under the cap:</strong> use your room.</li>
  <li><strong>Over the cap:</strong> use an exception if you are signing, or use matching or a TPE if you are trading.</li>
</ul>

<p>Everything else is a modifier on those two answers. The aprons do not add new rules so much as delete options from them, and the penalty currency shifts from money to tools to picks as you climb. A team keeps its own players freely at every level. What it loses, progressively, is the ability to go get anybody else.</p>

<div class="notes">
  <ol>
    <li id="fn1">Named for Larry Bird. The 1983 CBA introduced the league's first salary cap, and rather than make it hard, the owners and the union carved out exceptions, the central one being the right to exceed the cap to re-sign your own veteran. Formally it is the Veteran Free Agent Exception; Bird's name stuck because he re-signed with Boston as it came in. Whether the Celtics technically used it on Bird himself is disputed, since the cap took effect the following season, and Cedric Maxwell is often cited as the first real use. <a href="#fnref1">&#8617;</a></li>
    <li id="fn2">Sign-and-trade contracts were capped at four years by the 2011 CBA, matching what a team can offer using cap room. <a href="#fnref2">&#8617;</a></li>
  </ol>
</div>

<div class="footer">
  <p>Figures are 2025-26 and reset each summer; the structure does not. Rules per the 2023 CBA. Built from Larry Coon's Salary Cap FAQ, the Hoops Rumors glossary, NBA.com's annual cap release, and Spotrac for live payrolls. 2025 cap-space counts via ESPN. Boston figures and the Stevens quote via Boston.com, CBS Boston, FOX Sports, ESPN, and NBA.com's 2025 offseason trackers.</p>
  <p class="backlink"><a href="/">&larr; Mike Purewal</a> &nbsp;&middot;&nbsp; <a href="/writing/">All writing</a></p>
</div>

</body>
</html>]]></content><author><name></name></author><category term="sports-and-health" /><category term="strategy" /><category term="reference" /><summary type="html"><![CDATA[The NBA Salary Cap]]></summary></entry><entry><title type="html">Convexity</title><link href="/writing/2026/08/16/convexity.html" rel="alternate" type="text/html" title="Convexity" /><published>2026-08-16T00:00:00+00:00</published><updated>2026-08-16T00:00:00+00:00</updated><id>/writing/2026/08/16/convexity</id><content type="html" xml:base="/writing/2026/08/16/convexity.html"><![CDATA[<p>My drive to the train station on local roads takes ~10 minutes at 45 mph. Slowing to 35 costs me 3 mins, while
    speeding up to an unsafe 55 gives back less than 2. On longer distances, this asymmetry is more extreme. The human
    mind is built to think in straight lines, and that usually works, since any smooth function looks linear up close.
    That is the whole basis of a Taylor approximation. It breaks down when the second term is significant. In
    real life, this happens often. Here are some convex and concave function examples:</p>

<p><strong>Convex.</strong> Flat, then steep. </p>
<ul>
    <li>30 investments, 25 write-offs, and the returns come from one.</li>
    <li>A thousand falls from one foot leave you fine; one fall from a thousand feet does not.</li>
    <li>Many emergent behaviors: knowledge accumulates and does nothing for years until it produces the brilliant idea.</li>
</ul>

<p><strong>Concave.</strong> Steep, then flat. </p>
<ul>
    <li>The first $100k is life changing, the tenth doesn't change your lifestyle much (but still great!).</li>
    <li>Year 1 of practice makes you competent and is fun because you notice the change. Year 10 is grueling and makes
        you only slightly better than year 9.</li>
    <li>A reputation takes years to build and one event to destroy. The second event barely registers.</li>
</ul>

<p>I first heard the word <em>convexity</em> from a bond trader when I started my career, and it sounded like a cheat code: a bond with
    positive convexity gains more on a rally than it loses on an equivalent selloff. In 2013, I worked through Boyd's online lectures and his text
    <a href="https://web.stanford.edu/~boyd/cvxbook/">Convex Optimization</a>. Taleb's
    <a href="https://a.co/d/0gbb0qTU">Antifragile</a> is another great reference on the same idea.</p>

<p>Curvature distorts forecasting locally, and it breaks summary statistics globally. Jensen's inequality is the formal
    version: the average of a convex function is not the function of the average. Up 50% then down 50% averages zero but
    leaves you at 75. A river 4ft deep on average still drowns people.</p>

<p>Pressing the accelerator feels like it should pay, and we assume the payoff scales with the sensation. That instinct can be misleading.  It can also work in your favor: find places where the cost of
    many attempts are small and fixed, but the payoff is large. </p>]]></content><author><name></name></author><category term="writing" /><category term="decision-making" /><category term="mathematics" /><category term="risk" /><summary type="html"><![CDATA[My drive to the train station on local roads takes ~10 minutes at 45 mph. Slowing to 35 costs me 3 mins, while speeding up to an unsafe 55 gives back less than 2. On longer distances, this asymmetry is more extreme. The human mind is built to think in straight lines, and that usually works, since any smooth function looks linear up close. That is the whole basis of a Taylor approximation. It breaks down when the second term is significant. In real life, this happens often. Here are some convex and concave function examples:]]></summary></entry><entry><title type="html">From Hit Rate to Sharpe</title><link href="/2026/08/11/hit-rate-sharpe.html" rel="alternate" type="text/html" title="From Hit Rate to Sharpe" /><published>2026-08-11T15:00:00+00:00</published><updated>2026-08-11T15:00:00+00:00</updated><id>/2026/08/11/hit-rate-sharpe</id><content type="html" xml:base="/2026/08/11/hit-rate-sharpe.html"><![CDATA[<p>A hit rate carries no information about bet sizing, symmetry, or how often you trade. So turning one into a Sharpe ratio needs assumptions.</p>

<p><a href="/assets/files/hit-rate-sharpe/hit-rate-to-sharpe.html" target="_blank">From hit rate to Sharpe</a> starts with the simplest possible trade: equal size and independent. This gives a simple expression of a per-trade Sharpe of twice the hit-rate edge. Annualizing the per-trade Sharpe multiplies that by the square root of the number of bets, so 1,000 even-sized bets for reasonable hit rates is twice the edge times 32. Read backwards, the same relation says what accuracy a target Sharpe demands: a Sharpe-1.0 book needs a 51.6% hit rate for 100 bets.</p>

<p>Varying bet sizes, asymmetric payoffs, transaction costs, and correlation across simultaneous positions each pull you below that ceiling. Each is approximated in <a href="/assets/files/hit-rate-sharpe/discounts.html" target="_blank">The Discounts</a>. Taking rough numbers, like dispersion 1.0, costs 0.02 and correlation 0.2, a <a href="/2026/05/15/edges.html">54% hit rate</a> across 1,260 trades falls from an idealized annual Sharpe of <strong>2.85 to 1.12</strong>.</p>

<p>One interesting finding that I didn't appreciate was that the Sharpe ratio and the t-statistic turn out to be the same measurement: <em>t</em> = Sharpe × &radic;years, so over a single year they are the same number. Statistical significance and risk-adjusted return are two readings of one quantity. <a href="/assets/files/hit-rate-sharpe/trades-to-significance.html" target="_blank">From trade count to significance</a> works through the math. Roughly 1,400 trades to reach <em>t</em> = 3 at a 54% hit rate, and about 2,300 if you want an 80% chance of detecting an edge that is genuinely there.</p>

<p>To finish, <a href="/assets/files/hit-rate-sharpe/applet.html" target="_blank">an applet</a> that allows you to toggle all the variables (hit rate, trade count, etc.) to arrive at a Sharpe.</p>]]></content><author><name></name></author><category term="finance" /><category term="data-science" /><category term="tutorial" /><summary type="html"><![CDATA[A hit rate carries no information about bet sizing, symmetry, or how often you trade. So turning one into a Sharpe ratio needs assumptions.]]></summary></entry><entry><title type="html">Local Maxima</title><link href="/career/2026/05/27/local-maxima.html" rel="alternate" type="text/html" title="Local Maxima" /><published>2026-05-27T00:00:00+00:00</published><updated>2026-05-27T00:00:00+00:00</updated><id>/career/2026/05/27/local-maxima</id><content type="html" xml:base="/career/2026/05/27/local-maxima.html"><![CDATA[<p>Pursuing a career inside an organization means picking one of two modes. The <em>local max</em> strategy is simple: you don't need to outrun the bear, only the next person. You stay slightly ahead of your cohort, take the next promotion, never look like you're losing. It works. People build excellent careers this way. The cost is that your career becomes about the bear. Your ceiling is set by the peer group's floor; you're only as fast as you need to be. The skills you sharpen are the skills of staying ahead, which is not the same as the skills of going somewhere.</p>

<p>The alternative is the climber on an <em>infinite mountain</em>. There is no summit, only the next ridge. The mountain doesn't grade against other climbers and doesn't care who else is on it. This mode runs on what physicists call <em>simulated annealing</em>: you accept temporary descent to find a better route up. In an org, that means lateral moves, projects with no obvious payoff, stretches where you look like you're not winning. The local max runner cannot afford any of this.</p>

<p>The bear is real either way. Orgs are ruthless and if you're ambitious there's always a bear.  What path do you choose? A career spent outrunning the bear is not the same as a career spent climbing anything.</p>]]></content><author><name></name></author><category term="career" /><category term="career" /><category term="strategy" /><category term="ambition" /><category term="organizations" /><summary type="html"><![CDATA[Pursuing a career inside an organization means picking one of two modes. The local max strategy is simple: you don't need to outrun the bear, only the next person. You stay slightly ahead of your cohort, take the next promotion, never look like you're losing. It works. People build excellent careers this way. The cost is that your career becomes about the bear. Your ceiling is set by the peer group's floor; you're only as fast as you need to be. The skills you sharpen are the skills of staying ahead, which is not the same as the skills of going somewhere.]]></summary></entry><entry><title type="html">54%</title><link href="/2026/05/15/edges.html" rel="alternate" type="text/html" title="54%" /><published>2026-05-15T15:00:00+00:00</published><updated>2026-05-15T15:00:00+00:00</updated><id>/2026/05/15/edges</id><content type="html" xml:base="/2026/05/15/edges.html"><![CDATA[<p>Roger Federer once observed that he won about 80% of his matches, but only 54% of the points he played. Most fans would have guessed a much higher per-point edge for the greatest tennis player ever<sup>1</sup>. The actual edge was four percentage points above coinflip.  Across enough points, that compounded into the dominance everyone saw. Up close, in any given rally, the edge was invisible.</p>

<p>Professional investing skill has this shape. So does success in life more broadly. An information ratio of 1 corresponds to winning about 52.5% of trading days (<a href="/assets/files/IR_to_hit_rate_derivation.pdf" target="_blank">get derivation here</a>). The benchmark is essentially a coinflip on any individual day and the compounding is what makes it a career. The mechanics of how that compounding actually works are in <a href="/2019/10/26/improve-your-win-rate.html">Improve Your Win Rate</a>; what interests me here is why the underlying lean is so easy to miss in the first place.</p>

<p>The disconnect is between the <em>unit of skill</em> and the <em>unit of outcome</em>. Skill lives at the per-point, per-trade, per-decision level, where it shows up as a barely detectable lean against random. Outcome lives at the season, year, career level, where it shows up as visible dominance. This is fundamentally a <a href="/blog/2026/03/12/time-bars.html">sampling problem</a>: real skill is encoded at an interval far smaller than the one we actually pay attention to.</p>

<p>Discretionary traders rarely look obviously brilliant trade-by-trade. The good ones do not have a clear tell that separates them from average ones on any single decision, and the texture of their work, watched up close, often looks much like everyone else's. Expertise is <a href="/career/2026/03/02/stat-mech.html">emergent</a>, not accumulated.</p>

<p>You cannot trust your own perception of your own edge in real time, because your edge, if you have one, is too small to see directly while you have it. The people we dismiss as average might in fact have the same lean as the ones we celebrate, with the only difference being that their sample size has not yet caught up to make it visible to anyone, including themselves.</p>

<p>The work is to act on a lean you cannot feel, long enough for sample size to reveal whether one was ever there.</p>

<hr/>

<p><sup>1</sup><em>Thinking with Machines</em> by Vasant Dhar</p>]]></content><author><name></name></author><category term="data-science" /><category term="sports-and-health" /><category term="finance" /><category term="personal-development" /><summary type="html"><![CDATA[Roger Federer once observed that he won about 80% of his matches, but only 54% of the points he played. Most fans would have guessed a much higher per-point edge for the greatest tennis player ever1. The actual edge was four percentage points above coinflip. Across enough points, that compounded into the dominance everyone saw. Up close, in any given rally, the edge was invisible.]]></summary></entry><entry><title type="html">Optimism as Rigor</title><link href="/career/2026/04/14/optimism.html" rel="alternate" type="text/html" title="Optimism as Rigor" /><published>2026-04-14T00:00:00+00:00</published><updated>2026-04-14T00:00:00+00:00</updated><id>/career/2026/04/14/optimism</id><content type="html" xml:base="/career/2026/04/14/optimism.html"><![CDATA[<p>There's a phenomenon in sales performance that creates a gain of ~30%: how people narrate setbacks. Top performers treat failure as temporary, specific, and external: <em>this pitch didn't land with this audience on this day</em>. Underperformers treat it as permanent, pervasive, and internal: <em>I'm not good at this</em>. Are the optimists just deluding themselves?</p>

<p>I struggled with this. After a meeting that didn't go well, I'd leave convinced there was a deeper issue with my approach, something structural I was missing. This introspection felt like rigor. A mentor once stopped me mid-spiral and asked a simple question: <em>did anyone actually tell you it went poorly?</em> I'd run a full post-mortem on a failure that existed entirely in my own narration.</p>

<p>Pessimistic generalization feels like analysis because it's dressed in the language of thoroughness. I'd never accept that conclusion from a model or a colleague, so why accept it from myself? It's just as arbitrary as "wrong audience, wrong day."</p>

<p>The fix isn't to become an optimist. It's to apply the same standards to yourself that you already apply to everything else. When a meeting doesn't land, force the specificity: what exactly didn't work, for whom, and why. Not "what's wrong with my approach" but "what's the isolated variable." When you do this, the answers tend to be narrow and fixable, not existential.</p>

<p>That's the irony of the pessimist's explanatory style. True rigor applied to your own setbacks produces optimistic-sounding conclusions. The most disciplined response to a failed pitch is almost always closer to "wrong audience, wrong day" than <em>something is deeply wrong</em>. Stop over-extrapolating from N=1.</p>]]></content><author><name></name></author><category term="career" /><category term="career-advice" /><category term="strategy" /><summary type="html"><![CDATA[There's a phenomenon in sales performance that creates a gain of ~30%: how people narrate setbacks. Top performers treat failure as temporary, specific, and external: this pitch didn't land with this audience on this day. Underperformers treat it as permanent, pervasive, and internal: I'm not good at this. Are the optimists just deluding themselves?]]></summary></entry><entry><title type="html">Sampling Life</title><link href="/blog/2026/03/12/time-bars.html" rel="alternate" type="text/html" title="Sampling Life" /><published>2026-03-12T00:00:00+00:00</published><updated>2026-03-12T00:00:00+00:00</updated><id>/blog/2026/03/12/time-bars</id><content type="html" xml:base="/blog/2026/03/12/time-bars.html"><![CDATA[<p>In quantitative finance, your data sampling method changes what you see. Choose the wrong interval and real patterns disappear. I've been thinking lately about how the same problem applies to memory and how I reflect on my experiences.</p>

<p><strong>Calendar bars.</strong> The default sampling scheme for a life is <em>calendar bars</em>. Birthdays, new year's, decades, school years. They fire on schedule regardless of whether anything meaningful happened, which means transformative years are equivalent to years with relatively little change. The grid is even, but life experience isn't.</p>

<p><strong>Novelty bars.</strong> Memory seems to encode <em>novelty bars</em>. First apartment, first loss, that one semester. Ages 15 to 30 is a period of high novelty: new independence, new relationships, first experiences of basically everything. Your twenties feel like they lasted a decade because your brain was firing a bar often. Your late thirties feel like a weekend because of large intervals between novel experiences.</p>

<p><strong>Depth bars.</strong> The 100th weeknight dinner with family contains nominal novelty, which means memory essentially skips it. These are <em>depth bars</em>, periods where nothing new is happening but something is latently accumulating. Skill, familiarity, intuition, identity. They're invisible to the sampling scheme memory uses, but they're foundational in ways novelty bars never are.</p>

<p><strong>Emergence bars.</strong> Then there's a fourth type. A bar that fires when the cumulative deviation from expected behavior crosses a threshold. Nothing, nothing, nothing, and then a structural break: <em>emergence bars</em>. No single day of showing up at a job resembles the expertise that eventually materializes. No single conversation in a relationship resembles the closeness that eventually exists. Expertise is <a href="/career/2026/03/02/stat-mech.html">emergent</a>, not accumulated, and so the bar only fires in retrospect. When you try to locate it in memory you can't, because the moment of crossing wasn't a moment at all. It was a slow accumulation that only became apparent after the fact.</p>

<p>My reflection on life is written almost entirely in novelty bars, because those are the ones memory encodes. But the actual construction happened in depth bars and emergence bars that I can barely remember at all. The periods I skip over when reflecting are the ones that wrote it.</p>]]></content><author><name></name></author><category term="blog" /><category term="personal-development" /><summary type="html"><![CDATA[In quantitative finance, your data sampling method changes what you see. Choose the wrong interval and real patterns disappear. I've been thinking lately about how the same problem applies to memory and how I reflect on my experiences.]]></summary></entry><entry><title type="html">The Physics of Showing Up</title><link href="/career/2026/03/02/stat-mech.html" rel="alternate" type="text/html" title="The Physics of Showing Up" /><published>2026-03-02T00:00:00+00:00</published><updated>2026-03-02T00:00:00+00:00</updated><id>/career/2026/03/02/stat-mech</id><content type="html" xml:base="/career/2026/03/02/stat-mech.html"><![CDATA[<h3>Statistical Mechanics</h3>

<p>Of everything I encountered in physics, the most fascinating result was statistical mechanics' convergence of the microscopic and macroscopic. Two independent frameworks, one empirical and one theoretical, arrived at exactly the same mathematical structure. Entropy is a count of how many microscopic particle arrangements produce the same macroscopic observation. Entropy, temperature, and pressure <em>don't exist</em> at the micro level at all. They mathematically emerge when you aggregate across impossibly large numbers of particles, each one obeying simple rules and knowing nothing about the whole.</p>

<h3>Daily Work</h3>

<p>Individual days can feel arbitrary: another meeting, another bug fixed, another paper read, another problem half-solved. No single day contains your expertise the way no single molecule has temperature. You can't point at a Tuesday and say "that's where the knowledge lives." And yet the aggregate of thousands of those days produces <strong>judgment, expertise, taste</strong>, none of which can be traced to any specific moment.</p>

<h3>Compounding vs Emergence</h3>

<p>The metaphor I've always reached for is compound interest: small efforts add up. Compounding is accumulation: each unit contributes to a growing total which resembles the unit (just more of it). But statistical mechanics describes something different. The macroscopic property isn't a sum of the microscopic activity, it's a <em>qualitatively different</em> emergent behavior. Temperature isn't "a lot of kinetic energy added up." It's a new concept with no meaning at the particle level. In the same way, expertise is an emergent property of sustained exposure to hard problems.</p>

<p>This distinction matters because it changes the playbook. In the compounding framework, you should optimize each day, track your progress, and look for evidence that the total is growing. If daily work is a microstate in a <em>statistical ensemble</em>, none of that applies. You don't engineer temperature molecule by molecule. You set the conditions, sustain them, and the emergent property takes care of itself. The only thing that matters is that you keep showing up in environments where contact with hard problems is probable, and that you don't stop.</p>

<h3>Takeaway</h3>

<p>Emergence doesn't announce itself incrementally with a progress bar. It crystallizes when you don't expect it. One day, you know how to tackle a problem with a little more certainty because of the work put in. There can be long stretches where nothing seems to be happening, where the work feels meaningless and the days feel interchangeable. <strong>It is how emergence works.</strong> And the only way to achieve it is to keep showing up.</p>]]></content><author><name></name></author><category term="career" /><category term="career-advice" /><category term="strategy" /><summary type="html"><![CDATA[Statistical Mechanics]]></summary></entry></feed>