Part 1 · A single helper trait

Calculating the Gain in Predictive Accuracy From Adding a Helper Trait to a Baseline Model

This formula predicts the gain in predictive accuracy when adding a helper trait to a baseline predictive model, which is typically a polygenic score (PGS). This equation says that if we know a few characteristics that can be measured in empirical data, we can predict how much this will help us predict a target trait of interest.

This formula was derived in Zhang et al.

$$ R^2_{\mathrm{gain}} = \frac{ \left[\, \rho_g \sqrt{\frac{h_h^2}{h_t^2}} \left(h_t^2 - R^2_{\mathrm{base}}\right) + \rho_e \sqrt{\left(1 - h_h^2\right)\left(1 - h_t^2\right)} \,\right]^2 }{ R^2_{\mathrm{base}} \left[\, 1 - \rho_g^2 \, R^2_{\mathrm{base}} \, \frac{h_h^2}{h_t^2} \,\right] } $$

Equation (13), in the closed form derived in Additional file 1, Supplementary Methods 2.

How to use this page

  • Every section has the same five sliders. Move any one of them and every figure on the page redraws.
  • The Jump to a figure buttons in section 01 load a representative point from each figure in the paper.
  • ▶ Sweep \(h_t^2\) in section 04 animates the target heritability from 0.9 down to 0.1 and back, so you can watch which helper traits become the best ones.
  • Hover any curve or map for exact values; click a heat map to move the sliders to that point.
  • ◐ at the top right switches the page between light and dark appearance.

Contents

01

Where the correlation comes from

Two traits correlate through shared genes, shared environments, or both. How the correlation divides between those two sources decides whether a helper trait is worth measuring.

$$\rho_p \;=\; \underbrace{\rho_g\sqrt{h_t^2 h_h^2}}_{\text{genetic term}} \;+\; \underbrace{\rho_e\sqrt{(1-h_t^2)(1-h_h^2)}}_{\text{environmental term}}$$

Equation (7). The weights are functions of the two heritabilities: \(\rho_g\) is weighted by \(\sqrt{h_t^2h_h^2}\) and \(\rho_e\) by the two non-heritable fractions, \(\sqrt{(1-h_t^2)(1-h_h^2)}\).

Each trait's variance split into genetic and environmental parts, with the two correlations linking them

Each bar is one trait's variance, and the width of the blue block is that trait's heritability. The curved links are the two correlations: a thicker link means a stronger correlation, and a dashed link means a negative one.

The phenotypic correlation as a sum of a genetic and an environmental contribution

\(\rho_g\) is weighted by the two heritabilities and \(\rho_e\) by the two non-heritable fractions. Terms with opposite signs point in opposite directions, so they partly cancel.

Insight · from the paper

Because the two weights depend on the heritabilities, traits with the same phenotypic correlation can be built from completely different mixes of shared genes and shared environment — and \(\rho_p\) alone cannot tell them apart. Read backwards, this equation also gives a route to \(\rho_e\), the environmental correlation, which is otherwise hard to observe directly.

02

Response to each parameter

Each panel varies one input and holds the other four where you left them, so every panel redraws whenever you move a slider.

$$ R^2_{\mathrm{gain}} = \frac{ \left[\, \rho_g \sqrt{\frac{h_h^2}{h_t^2}} \left(h_t^2 - R^2_{\mathrm{base}}\right) + \rho_e \sqrt{\left(1 - h_h^2\right)\left(1 - h_t^2\right)} \,\right]^2 }{ R^2_{\mathrm{base}} \left[\, 1 - \rho_g^2 \, R^2_{\mathrm{base}} \, \frac{h_h^2}{h_t^2} \,\right] } $$

Equation (13), in closed form. Each panel below is a slice through it.

\(R^2_{\mathrm{gain}}\) current value outside the model
as a table
Insight · from the paper

Once the other correlation is nonzero, the response is not symmetric about zero. Same-sign genetic and environmental correlations reinforce one another, and opposite signs cancel. So when \(\rho_e\) is positive, \(\rho_g = +0.5\) beats \(\rho_g = -0.5\); when \(\rho_e\) is negative, that ordering reverses. With the other correlation at exactly zero the response is symmetric. Over the ranges explored in Figure 2 of the paper, helper traits improved accuracy by up to about 70%.

03

What drives the gain

The Methods split the absolute gain into three terms, in the expression that follows Equation (12). Dividing by the baseline turns it into the relative gain of Equation (13). Any term near zero collapses the result.

$$\Delta R^2 \;=\; \left(1 - R^2_{\mathrm{base}}\right)\;\times\; r^2_{th}\;\times\;\left(1 - \mathcal{R}\right)$$
$$R^2_{\mathrm{gain}} \;=\; \frac{\Delta R^2}{R^2_{\mathrm{base}}}$$

Insight · from the paper

A helper trait can correlate strongly with the target and still be worthless if the baseline already captures that correlation. The paper's empirical selection tackles a related but different problem: redundancy among the candidate helper traits, filtered greedily on their pairwise correlations. Redundancy with the polygenic score itself, the \(\mathcal{R}\) of the theory, does not enter that selection.

04

The whole landscape

These are Figures 3 and 4 of the paper, recomputed live: at \(h_t^2 = 0.9\) the two panels reproduce Figure 3, and at \(h_t^2 = 0.1\) they reproduce Figure 4. Click ▶ Sweep \(h_t^2\) and the target heritability travels from \(h_t^2 = 0.9\) down to \(h_t^2 = 0.1\) and back, carrying both panels from one figure to the other. While it runs, the sweep also drives \(R^2_{\mathrm{base}}\), holding \(\alpha = 0.1\) as both figures do; the value you set is restored when you stop. The two panels respond very differently. Panel A, with the environmental correlation on its vertical axis, transforms: as the target becomes less heritable there is more environmental variance for \(\rho_e\) to work with, the largest gains grow by roughly ten to fifty times, and the darkest cells move left, to low helper heritability. Panel B, with the genetic correlation on its vertical axis, changes far less, because the genetic channel is capped by how much heritable variance the target has to share.

0

as a table
Insight · from the paper

For a highly heritable target, the best helper traits are heritable too. For a weakly heritable target the ranking flips: low-heritability helper traits often give the largest gains, because what they add is environmental signal a polygenic score cannot contain, and for such a target the gains persist even against an already-accurate baseline.

05

Same correlation, different values

Every helper trait below has the same \(\rho_p\) with the target, and only the source of that correlation differs. Set \(\rho_p\) on the left, then drag \(\rho_g\): \(\rho_e\) adjusts to keep \(\rho_p\) fixed, so you move along one family of equally correlated helper traits.

Gain across helper traits sharing the same phenotypic correlation

Insight · from the paper

Two helper traits with the same \(\rho_p\) can differ substantially in what they contribute. The genetic and environmental architecture underneath is what decides.