Skip to content

Latest commit

 

History

History
899 lines (760 loc) · 36.2 KB

File metadata and controls

899 lines (760 loc) · 36.2 KB

Chart Types — Statistical model output (§8-§12, §17-§19, §32-§33)

Charts for displaying statistical model results: forest plots, path / SEM diagrams (including IV), mediation (triangular / serial / parallel), moderation and simple slopes, factor / class models, growth curves, specification curves, and causal DAGs.

All HEX values, figure sizes, and helper signatures are defined in scripts/ssci_style.py (SSOT). This file references them by name only.

See chart-type-guide.md for the full catalog index. For descriptive distribution and bivariate charts, see chart-types-core.md.


8. Forest Plot

When to use: Meta-analysis effect size summary; subgroup analysis; also increasingly used to display regression coefficients (coefplot — see §8.1 for the single-model variant and §26 for the multi-model facet variant).

Key parameters:

  • Size: FIGURE_SIZES['forest_plot']
  • Point: circle/square, size proportional to study weight
  • CI line: horizontal, NEUTRAL['error_bar'], capsize=0
  • Null line: vertical dashed at effect = 0 (or OR = 1)

Core code pattern:

from matplotlib.patches import Polygon

focal = get_palette(1)[0]
for i, (name, es, ci_lo, ci_hi, weight) in enumerate(studies):
    ax.errorbar(es, i, xerr=[[es - ci_lo], [ci_hi - es]],
                fmt='o', color=focal, markersize=3 + weight * 8,
                ecolor=NEUTRAL['error_bar'], elinewidth=1.0, capsize=0)
    ax.text(-margin, i, name, va='center', ha='right', fontsize=8)
    ax.text(right_margin, i, f'{es:.2f} [{ci_lo:.2f}, {ci_hi:.2f}]',
            va='center', ha='left', fontsize=8)
ax.axvline(x=0, color=NEUTRAL['reference'], linestyle='--', linewidth=0.8)
# Diamond for overall effect
diamond = Polygon([[es_overall-ci_w, y_pos], [es_overall, y_pos+0.3],
                   [es_overall+ci_w, y_pos], [es_overall, y_pos-0.3]],
                  closed=True, facecolor=focal)
ax.add_patch(diamond)

APA notes:

  • Left column: study names (left-aligned); right column: effect [CI] numeric
  • Report heterogeneity (I², Q, τ²) in Note — I and Q italic (Latin); Greek τ not italic
  • Diamond width = CI of overall effect
  • Subgroups: horizontal separator line + subgroup subtotal diamond
  • CI uses brackets: 95% CI [0.25, 0.75]

Common mistakes:

  • Marker size not varying by weight (all same size loses precision context)
  • Missing null reference line
  • CI notation wrong — APA uses brackets: 95% CI [0.25, 0.75], not (...)

8.1 Coefficient / Dot-and-Whisker Variant

When to use: Multiple regression coefficients with CI; political science / economics / management. Visually identical layout to §8 forest, but semantically different: coefficients (sign-free range) instead of effect sizes (positive bound). For a multi-model dashboard where Model 1 / Model 2 / Robust are compared side-by-side, see §26.

Key differences from §8:

  • Marker size constant (no weight encoding) unless N varies dramatically
  • Null reference line at 0 (not at OR = 1 / d = 0)
  • Right column: b [95% CI] (no overall pooled diamond at bottom)
  • Category groupings: horizontal subgroup separators per category (e.g., HR practices / firm controls / interactions)

Core code pattern:

fig, ax = plt.subplots(figsize=FIGURE_SIZES['forest_plot'])
focal, ref = get_emphasis_pair()
for i, (coef_name, b, lo, hi, sig) in enumerate(coefs):
    color = focal if sig else ref
    ax.errorbar(b, i, xerr=[[b - lo], [hi - b]],
                fmt='o', color=color, markersize=5,
                ecolor=color, elinewidth=1.0, capsize=0)
    ax.text(-margin, i, coef_name,
            va='center', ha='right', fontsize=8)
    ax.text(right_margin, i,
            f'{b:.2f} [{lo:.2f}, {hi:.2f}]',
            va='center', ha='left', fontsize=8)
add_reference_line(ax, value=0, orientation='v')

For multi-model facet layout (e.g. Model 1 / Model 2 / Robust), see §26 in chart-types-causal-econ.md and multi-panel.md compose_grid() / small_multiples().

APA notes:

  • b italic (Latin coefficient); β (standardized) not italic
  • State whether SE / CI / robust SE in Note
  • Sort by substantive grouping, not magnitude

Representative paper: Hainmueller & Hopkins (2014, APSR) conjoint analysis; Solt & Hu (2024) dotwhisker R package.

Common mistakes:

  • Using non-zero null line for raw coefficients (only OR/HR get OR = 1; raw b gets 0)
  • Forgetting per-category separator brackets when there are 3+ thematic groups
  • Sorting by magnitude rather than substantive grouping (kills comparison across panels in §26)

9. Path Diagram / SEM

When to use: Structural equation modeling; CFA measurement models; direct/indirect effect paths.

Key parameters:

  • Size: FIGURE_SIZES['path_model']
  • Observed variables: FancyBboxPatch (rounded rectangle), white fill
  • Latent variables: Ellipse, white fill
  • Significant path: solid line 1.2--1.5 pt; non-significant: dashed 1.0 pt, color NEUTRAL['ns_path']

Core code pattern:

from matplotlib.patches import FancyBboxPatch, FancyArrowPatch, Ellipse

# Latent variable (ellipse)
ellipse = Ellipse((x, y), width=1.2, height=0.6, fill=True,
                  facecolor='white', edgecolor='black', linewidth=1.0)
ax.add_patch(ellipse)
# Observed variable (rounded rect)
box = FancyBboxPatch((x-0.5, y-0.25), 1.0, 0.5,
                      boxstyle='round,pad=0.05', facecolor='white',
                      edgecolor='black', linewidth=1.0)
ax.add_patch(box)
# Path arrow with coefficient
arrow = FancyArrowPatch((x1, y1), (x2, y2),
                         arrowstyle='->', mutation_scale=12,
                         linewidth=1.5 if sig else 1.0,
                         linestyle='-' if sig else '--',
                         color='black' if sig else NEUTRAL['ns_path'])
ax.add_patch(arrow)
ax.text(mid_x, mid_y, r'$\beta$ = .45***', fontsize=8, ha='center')

APA notes:

  • Path coefficients: standardized, italic label (a, b, c' are algebraic variables -- italic)
  • Greek letters (β, χ, τ) are not italic per APA 7 Table 6.5
  • Report fit indices in Note: χ², CFI, TLI, RMSEA, SRMR
  • Layout: X (left) → M (center) → Y (right); covariates at top or bottom

Common mistakes:

  • Italicizing Greek letters (β, η) — only Latin letters are italic
  • Forgetting error/residual terms for endogenous variables
  • Curved double-headed arrow missing for correlated errors
  • Not reporting model fit indices

9.1 Instrumental Variable (IV) Diagram

When to use: Display an IV identification strategy — three roles (Z instrument, X endogenous, Y outcome) with explicit exclusion restriction. Standard in economics / political science / epidemiology where the identification strategy itself is the contribution.

This subsection is also referenced as §38 IV Diagram under the causal-inference family — the content is one and the same; the cross- reference acknowledges that the chart belongs structurally to the path / SEM family (shared FancyArrowPatch code mode) but is used disproportionately in causal-econ work. See chart-types-causal-econ.md.

Key differences from §9:

  • Three nodes in horizontal chain (Z → X → Y), no latent variables in the standard layout (optional unmeasured U shown as ellipse)
  • Exclusion restriction: a dashed curved arrow from Z to Y with an "Excluded by design" label
  • First-stage strength annotation: F statistic near Z → X arrow
  • Reduced-form / 2SLS coefficient: inline annotation near Y endpoint

Core code pattern:

from matplotlib.patches import FancyBboxPatch, FancyArrowPatch

fig, ax = plt.subplots(figsize=FIGURE_SIZES['path_model'])
ax.set_xlim(0, 1); ax.set_ylim(0, 1); ax.axis('off')

nodes = {'Z': (0.15, 0.5), 'X': (0.5, 0.5), 'Y': (0.85, 0.5)}
for name, (nx, ny) in nodes.items():
    box = FancyBboxPatch((nx - 0.06, ny - 0.05), 0.12, 0.10,
                          boxstyle='round,pad=0.02', facecolor='white',
                          edgecolor='black', linewidth=1.0)
    ax.add_patch(box)
    ax.text(nx, ny, name, ha='center', va='center',
            fontsize=10, fontweight='bold')

# Strong-instrument first-stage arrow Z -> X
ax.add_patch(FancyArrowPatch(nodes['Z'], nodes['X'],
                             arrowstyle='->', mutation_scale=14,
                             linewidth=1.6, color='black'))
ax.text(0.32, 0.55, r'$F$ = 28.4, $p$ < .001',
        fontsize=8, ha='center')

# Second-stage causal arrow X -> Y
ax.add_patch(FancyArrowPatch(nodes['X'], nodes['Y'],
                             arrowstyle='->', mutation_scale=14,
                             linewidth=1.6, color='black'))
ax.text(0.68, 0.55, r"$\beta_{\mathrm{IV}}$ = .38",
        fontsize=8, ha='center')

# Excluded path Z -> Y (dashed, curved below the chain)
ax.add_patch(FancyArrowPatch(nodes['Z'], nodes['Y'],
                             arrowstyle='->', mutation_scale=10,
                             linewidth=1.0, linestyle='--',
                             color=NEUTRAL['ns_path'],
                             connectionstyle='arc3,rad=-0.35'))
ax.text(0.5, 0.22, 'Excluded by design',
        fontsize=7, color=NEUTRAL['ns_path'],
        ha='center', style='italic')

APA notes:

  • F and p italic (Latin); Greek β (lowercase) not italic
  • Report first-stage F (Stock-Yogo threshold), 2SLS coefficient, n
  • Optional: an unmeasured confounder U drawn as a latent (Ellipse) below the X-Y line, with arrows into both X and Y, to make the identification argument visual

Common mistakes:

  • Omitting the exclusion-restriction visualization (the whole point of IV)
  • Confusing first-stage / reduced-form / 2SLS coefficients
  • Strong-instrument check missing (F > 10 rule of thumb; Stock-Yogo critical values for the relevant n and weak-IV bias tolerance)

10. Mediation Diagram

When to use: Mediation pathway display; indirect effect decomposition (X → M → Y).

Key parameters:

  • Size: FIGURE_SIZES['mediation']
  • Layout: triangular -- X (left), M (top), Y (right)
  • Path labels: a (X → M), b (M → Y), c' (X → Y direct)

Core code pattern:

from matplotlib.patches import FancyBboxPatch, FancyArrowPatch

# Three nodes in triangle layout
nodes = {'X': (0.15, 0.3), 'M': (0.5, 0.8), 'Y': (0.85, 0.3)}
for name, (nx, ny) in nodes.items():
    box = FancyBboxPatch((nx-0.08, ny-0.04), 0.16, 0.08,
                          boxstyle='round,pad=0.02', facecolor='white',
                          edgecolor='black', linewidth=1.0,
                          transform=ax.transAxes)
    ax.add_patch(box)
    ax.text(nx, ny, var_labels[name], transform=ax.transAxes,
            ha='center', va='center', fontsize=9)
# Paths a, b, c'
for (start, end, coef, sig, label) in paths:
    arrow = FancyArrowPatch(start, end, arrowstyle='->',
                             mutation_scale=12,
                             linewidth=1.5 if sig else 1.0,
                             linestyle='-' if sig else '--',
                             color='black' if sig else NEUTRAL['ns_path'],
                             transform=ax.transAxes)
    ax.add_patch(arrow)

APA notes:

  • Report indirect effect ab and its CI (bootstrap) in Note
  • Total effect c = c' + ab
  • Path labels (a, b, c') italic -- they are algebraic variables
  • Report model fit indices in Note; see §9 APA notes for the standard fit index template

Common mistakes:

  • Forgetting the direct path c' (X → Y)
  • Not reporting bootstrap CI for indirect effect
  • Confusing unstandardized and standardized coefficients

Layout convention. This module uses the triangular layout (M at apex, X-Y at base) per Baron & Kenny (1986). The direct path c' is a straight horizontal arrow at the base. If you adapt this template to a linear chain (X-M-Y on one row, e.g. §10.1 serial), c' is curved as an arc below the row to avoid passing through M. The curvature is geometric (collision avoidance), not semantic — per SEM convention, curved single-headed arrows remain regression paths; double-headed curved arrows denote covariances. The indirect effect ab is reported as a numerical value with bootstrap CI in the figure Note, never as a separate arrow. See apa-figure-standards.md §11 ("Mediation Diagram Conventions") for full SEM notation rules.

10.1 Serial Mediation (X → M1 → M2 → Y)

Layout: Horizontal chain -- X (far left), M1 (center-left), M2 (center-right), Y (far right). Place nodes at the same vertical center, with paths flowing left to right.

# Four nodes in horizontal chain layout
nodes = {'X': (0.10, 0.5), 'M1': (0.37, 0.5),
         'M2': (0.63, 0.5), 'Y': (0.90, 0.5)}
# Paths: a1 (X->M1), d21 (M1->M2), b2 (M2->Y), c' (X->Y, curved below)
paths = [
    (nodes['X'], nodes['M1'], a1_coef, a1_sig, '$a_1$'),
    (nodes['M1'], nodes['M2'], d21_coef, d21_sig, '$d_{21}$'),
    (nodes['M2'], nodes['Y'], b2_coef, b2_sig, '$b_2$'),
]
# Direct path c' drawn as a curved arrow below the chain
from matplotlib.patches import FancyArrowPatch
direct = FancyArrowPatch(nodes['X'], nodes['Y'],
                          connectionstyle='arc3,rad=-0.3',
                          arrowstyle='->', mutation_scale=12,
                          linewidth=1.5 if cp_sig else 1.0,
                          linestyle='-' if cp_sig else '--',
                          color='black' if cp_sig else NEUTRAL['ns_path'],
                          transform=ax.transAxes)
ax.add_patch(direct)
# Label c' at the midpoint of the curve
ax.text(0.5, 0.25, f"$c'$ = {cp_coef:.2f}",
        transform=ax.transAxes, ha='center', fontsize=8)

Note template: "Indirect effect through M1 and M2: a₁d₂₁b₂ = .XX, 95% CI [.XX, .XX]."

The curvature of c' in this layout is a geometric collision-avoidance choice — see the Layout convention paragraph in §10 above and apa-figure-standards.md §11 for full SEM notation justification.

10.2 Parallel Mediation (X → M1/M2 → Y)

Layout: Diamond -- X (left), M1 (top), M2 (bottom), Y (right). Both mediators share the same horizontal center but are vertically separated.

# Four nodes in diamond layout
nodes = {'X': (0.10, 0.5), 'M1': (0.5, 0.82),
         'M2': (0.5, 0.18), 'Y': (0.90, 0.5)}
# Paths: a1 (X->M1), a2 (X->M2), b1 (M1->Y), b2 (M2->Y),
#        c' (X->Y, through center)
paths = [
    (nodes['X'], nodes['M1'], a1_coef, a1_sig, '$a_1$'),
    (nodes['X'], nodes['M2'], a2_coef, a2_sig, '$a_2$'),
    (nodes['M1'], nodes['Y'], b1_coef, b1_sig, '$b_1$'),
    (nodes['M2'], nodes['Y'], b2_coef, b2_sig, '$b_2$'),
]
# Direct path c' as a straight horizontal arrow through center
direct = FancyArrowPatch(nodes['X'], nodes['Y'],
                          arrowstyle='->', mutation_scale=12,
                          linewidth=1.5 if cp_sig else 1.0,
                          linestyle='-' if cp_sig else '--',
                          color='black' if cp_sig else NEUTRAL['ns_path'],
                          transform=ax.transAxes)
ax.add_patch(direct)
ax.text(0.5, 0.55, f"$c'$ = {cp_coef:.2f}",
        transform=ax.transAxes, ha='center', fontsize=8)

Note template: "Indirect effect through M1: a₁b₁ = .XX, 95% CI [.XX, .XX]. Indirect effect through M2: a₂b₂ = .XX, 95% CI [.XX, .XX]. Report each indirect effect and its bootstrap CI separately."

The straight c' in this layout follows §10 main convention — see apa-figure-standards.md §11 for the full SEM rules.


11. Interaction / Moderation Plot

When to use: Visualizing moderation effects; interaction decomposition in ANOVA or regression.

Key parameters:

  • Size: FIGURE_SIZES['interaction']
  • Lines: one per moderator level; use LINE_STYLES[i] + MARKERS[i]
  • Moderator levels: typically −1SD, Mean, +1SD for continuous W

Core code pattern:

colors = get_palette(3)  # low, mean, high moderator
x_range = np.linspace(x_min, x_max, 100)
for i, (w_val, label) in enumerate(moderator_levels):
    # Y = b0 + b1*X + b2*W + b3*X*W
    y_pred = b0 + b1 * x_range + b2 * w_val + b3 * x_range * w_val
    ax.plot(x_range, y_pred, color=colors[i], linestyle=LINE_STYLES[i],
            marker=MARKERS[i], markevery=20, label=label)
ax.legend(title='Moderator', frameon=False)

APA notes:

  • Label moderator levels clearly: "Low (−1SD)", "Mean", "High (+1SD)"
  • Do not extrapolate beyond observed data range
  • State the regression equation or model specification in text or Note
  • SD italic (Latin statistical abbreviation)

Common mistakes:

  • Lines extending beyond observed data range
  • Unlabeled moderator levels
  • Only showing 2 levels when 3 (low/mean/high) are standard

11.1 Categorical Moderator

When the moderator is categorical (e.g., gender, treatment group) rather than continuous, plot one regression line per group instead of evaluating at ±1SD.

Key differences from continuous moderator:

  • Number of lines = number of groups (not the standard 3 levels)
  • Legend labels use group names directly ("Male", "Female") instead of "Low (−1SD)", "High (+1SD)"
  • Each group's regression line is fitted separately or derived from the interaction term
colors = get_palette(n_groups)
x_range = np.linspace(x_min, x_max, 100)
for i, (group_label, intercept, slope) in enumerate(group_slopes):
    y_pred = intercept + slope * x_range
    ax.plot(x_range, y_pred, color=colors[i], linestyle=LINE_STYLES[i],
            marker=MARKERS[i], markevery=20, label=group_label)
    # Overlay raw data points for each group
    mask = group_var == group_label
    ax.scatter(x[mask], y[mask], color=colors[i], alpha=0.3, s=15,
               marker=MARKERS[i], edgecolors='none')
ax.legend(title='Group', frameon=False)

When to use which:

  • Continuous W (e.g., self-esteem score): evaluate at −1SD, Mean, +1SD per the §11 main pattern
  • Categorical W (e.g., gender, condition): plot one line per group per §11.1
  • For slope significance annotations with either type, see §12 (Simple Slopes); for continuous marginal effect over W, see §11.2 below

11.2 Continuous Marginal Effect Plot

When to use: Continuous moderator W; show the marginal effect of X on Y as a function of W, with pointwise CI bands. Standard in political science (Brambor-Clark-Golder 2006) and economics; complements §11 which shows three discrete levels.

Key differences from §11:

  • X-axis = continuous W (not the focal predictor)
  • Y-axis = ∂Y/∂X (marginal effect, can be negative)
  • Single line + CI band (not three lines per W level)
  • Reference line at marginal effect = 0 (when CI crosses, the effect transitions to/from significance)

Core code pattern:

fig, ax = plt.subplots(figsize=FIGURE_SIZES['interaction'])
w_grid = np.linspace(w_min, w_max, 100)
me = b1 + b3 * w_grid           # marginal effect of X on Y at each W
se_me = np.sqrt(var_b1 + (w_grid**2) * var_b3
                + 2 * w_grid * cov_b1_b3)
ci_lo = me - 1.96 * se_me
ci_hi = me + 1.96 * se_me

focal = get_palette(1)[0]
ax.plot(w_grid, me, color=focal, lw=1.6)
ax.fill_between(w_grid, ci_lo, ci_hi, color=focal, alpha=0.18)
add_reference_line(ax, value=0, orientation='h',
                   label='No effect')
ax.set_xlabel(r'Moderator $W$')
ax.set_ylabel(r'Marginal effect $\partial Y / \partial X$')

APA notes:

  • Report b₁, b₃, their variances and covariance from the regression model
  • Mark where CI crosses zero with vertical dashed lines (Johnson-Neyman points; see §15.5 in chart-types-quick-reference.md)
  • W italic (Latin variable name)

Common mistakes:

  • Treating constant-W marginal effect as the X-on-Y main effect (it is not, unless b₃ = 0)
  • Forgetting to include cov(b1, b3) in the SE computation (gives the wrong CI width)
  • Extrapolating beyond observed W range

12. Simple Slopes Plot

When to use: Detailed moderation display with slope significance; complements the interaction plot with annotated slope values.

Key parameters:

  • Size: FIGURE_SIZES['interaction']
  • Same line/marker encoding as interaction plot
  • CI bands: fill_between, alpha=0.15

Core code pattern:

colors = get_palette(3)
# Each tuple: (w_value, intercept_at_w, slope, se, p, label)
for i, (w_val, intercept, slope, se, p, label) in enumerate(simple_slopes):
    y_pred = intercept + slope * x_range
    # Simplified CI for visualization only — see Common mistakes
    ci_hi = y_pred + 1.96 * se * x_range
    ci_lo = y_pred - 1.96 * se * x_range
    ax.plot(x_range, y_pred, color=colors[i],
            linestyle=LINE_STYLES[i], label=label)
    ax.fill_between(x_range, ci_lo, ci_hi, color=colors[i], alpha=0.15)
    # Annotate slope
    stars = p_to_stars(p)
    ax.text(x_range[-1], y_pred[-1],
            f'  $b$ = {slope:.2f}{stars}', fontsize=7,
            color=colors[i], va='center')

APA notes:

  • Annotate each slope (b, italic) with significance stars
  • Report simple slope tests in Note or text
  • 95% CI bands around each regression line are increasingly expected

Common mistakes:

  • Missing slope value annotations (plot alone is not enough)
  • CI bands for different slopes overlapping without distinguishable fill pattern
  • Extrapolating lines beyond observed X range
  • The simplified CI shown above is for visualization only; the exact CI requires the full variance-covariance matrix — use lavaan / statsmodels for inferential reporting

17. Scree Plot + Parallel Analysis

When to use: Determine factor retention in EFA / PCA. Show eigenvalues vs factor number; overlay parallel analysis (eigenvalues from random data simulations); retain factors where observed exceeds the simulated 95th percentile (Horn, 1965; Glorfeld, 1995).

Key parameters:

  • Size: FIGURE_SIZES['scree']
  • Palette: get_emphasis_pair() — observed focal vs simulated reference
  • Markers: 'o' (observed) connected by line; 's' (simulated) with dashed line
  • Optional: Kaiser horizontal at eigenvalue = 1 (only strictly valid for PCA on a correlation matrix)

Core code pattern:

fig, ax = plt.subplots(figsize=FIGURE_SIZES['scree'])
focal, ref = get_emphasis_pair()
factor_n = np.arange(1, len(eigenvalues) + 1)

ax.plot(factor_n, eigenvalues, 'o-', color=focal, lw=1.4, ms=6,
        label='Observed eigenvalue')
ax.plot(factor_n, sim_95th, 's--', color=ref, lw=1.0, ms=5,
        label='Parallel analysis 95th percentile')
# Optional Kaiser horizontal at eigenvalue = 1
add_reference_line(ax, 1.0, 'h', label='Kaiser (PCA)')
# Find retention point
retain_k = int(np.argmax(eigenvalues < sim_95th))
ax.axvline(retain_k - 0.5, color=NEUTRAL['axis'],
           linestyle=':', lw=0.6)
ax.annotate(f'Retain {retain_k} factors',
            xy=(retain_k - 0.5, ax.get_ylim()[1] * 0.9),
            xytext=(5, 0), textcoords='offset points', fontsize=7)
ax.set_xlabel('Factor number'); ax.set_ylabel('Eigenvalue')
ax.legend(frameon=False)

APA notes:

  • State n iterations of parallel analysis (typical 500-1000)
  • "Eigenvalue" is a Latin word, not a symbol — set in roman, not italic
  • For PCA: report variance explained for retained factors
  • For EFA: report fit indices for the retained-factor model

Common mistakes:

  • Using Kaiser-only with EFA (use parallel analysis or MAP)
  • Too few parallel-analysis iterations (loss of precision)
  • Not stating extraction method (PA / ML / PAF / OLS)
  • Confusing scree-style eigenvalue plots for PCA vs EFA — both look similar but the interpretation differs

Representative software: psych R package fa.parallel(); Glorfeld (1995) eigenvalue percentile correction.


18. Profile Plot (LPA/LCA)

When to use: Display latent profile analysis (LPA) or latent class analysis (LCA) results: each profile / class is one line; x = indicators; y = standardized means (z-scores); profile labels carry class proportions. Standard in personality typologies (Lanza & Rhoades, 2013) and developmental subgrouping.

Key parameters:

  • Size: FIGURE_SIZES['profile']
  • Palette: get_palette(n_profiles)
  • Lines: one per profile; LINE_STYLES[i] + MARKERS[i]
  • Reference horizontal at z = 0 (overall sample mean)
  • Profile labels with proportions: "Profile 1 (27%)"

Core code pattern:

fig, ax = plt.subplots(figsize=FIGURE_SIZES['profile'])
colors = get_palette(n_profiles)
for i, (mean_vector, label, prop) in enumerate(profiles):
    ax.plot(range(n_indicators), mean_vector,
            color=colors[i], linestyle=LINE_STYLES[i],
            marker=MARKERS[i], markersize=6, lw=1.5,
            label=f'{label} ({prop * 100:.0f}%)')
add_reference_line(ax, 0, 'h', label='Overall sample mean')
ax.set_xticks(range(n_indicators))
ax.set_xticklabels(indicator_names, rotation=30, ha='right')
ax.set_ylabel('Standardized indicator score ($z$)')
ax.legend(title='Profile (size)', frameon=False, loc='best')

APA notes:

  • State extraction method (LPA = continuous indicators; LCA = categorical)
  • Report fit indices in caption (BIC, aBIC, entropy, LMR-LRT)
  • Profile proportions sum to 1.0
  • Convention: order profiles by size (largest to smallest)

Common mistakes:

  • Using raw means (not standardized; loses cross-indicator comparison)
  • Profile labels without sample proportion
  • Connecting lines across non-ordinal indicators (interpretation risk — a line implies an ordering)
  • Forgetting fit indices in the figure Note

Representative software: tidyLPA R package; Mplus output exported to ggplot for visualization.


19. Growth Curve Plot

When to use: Longitudinal trajectories — individual curves in light gray (transparent "spaghetti") plus group-mean model-implied curves with CI bands. Standard in education (growth in literacy / math achievement) and developmental psychology (motor / cognitive milestones).

Key parameters:

  • Size: FIGURE_SIZES['growth']
  • Palette: get_palette(n_groups) — group means; individual trajectories in transparent versions of each group's color
  • Individual lines: thin (lw=0.3-0.4), low alpha (0.15-0.25)
  • Group means: bold (lw=2.0), saturated colors
  • CI bands: alpha=0.18, matching the group mean color
  • X-axis: substantive time scale (age in years preferred over wave number) with integer ticks at measurement points

Core code pattern:

fig, ax = plt.subplots(figsize=FIGURE_SIZES['growth'])
colors = get_palette(n_groups)

# Individual trajectories (background, all groups)
for traj in individual_data:
    group_id = traj['group']
    color = colors[group_id]
    ax.plot(traj['time'], traj['y'], color=color,
            lw=0.3, alpha=0.15, zorder=1)

# Group means with CI bands
for g in range(n_groups):
    mean = group_means[g]
    lo, hi = group_ci_lo[g], group_ci_hi[g]
    ax.plot(t_grid, mean, color=colors[g], lw=2.0,
            marker='o', ms=5, zorder=3,
            label=f'{group_labels[g]} (N={n_per_group[g]})')
    ax.fill_between(t_grid, lo, hi, color=colors[g],
                    alpha=0.18, zorder=2)
ax.set_xlabel('Time / wave'); ax.set_ylabel('Outcome')
ax.legend(frameon=False, loc='best')

APA notes:

  • State the growth model (linear / quadratic; random intercept / slope)
  • Variance components: σ²_intercept, σ²_slope, σ²_residual reported in the Note
  • Greek σ not italic (APA 7 Table 6.5)
  • Time should be on a substantive scale where possible (age in years preferred over wave number; participant age makes findings comparable across studies)

Common mistakes:

  • Showing individual lines without group means (chaos — the eye loses the structure)
  • Showing group means without individual lines (loses heterogeneity)
  • Too many individuals (>200) creates a hairball — sample to a representative subset or switch to a heatmap-of-trajectories
  • Forgetting variance components in the caption (without them, the figure cannot be reproduced or compared)

Representative software: lme4 / nlme R packages; AERJ / Developmental Psychology standard figure templates.


32. Specification Curve / Multiverse Plot

When to use: Test the robustness of an estimate to N analytic decisions (e.g., outlier exclusion choice, covariate set, transformation). Plot ordered effects + a descriptor matrix indicating which decisions produced each effect. The methodological-transparency gold standard for replication and robustness reports (Botvinik-Nezer et al., 2020 NARPS; Simonsohn et al., 2020 NHB; Steegen et al., 2016).

This is a two-panel composite chart (top: spec curve; bottom: descriptor matrix). Use compose_grid('A;B', height_ratios=[1, 1.4]) for layout.

Key parameters:

  • Size: FIGURE_SIZES['spec_curve']
  • Top panel: scatter (one point per specification, sorted by effect) + optional CI bars via errorbar
  • Bottom panel: binary descriptor matrix (rows = analytic choices, columns = specifications)
  • Reference line: vertical at effect = 0 (or the relevant null)
  • Color: sig vs ns (95% CI excludes 0) via get_emphasis_pair()

Core code pattern:

fig, axes = compose_grid([['A'], ['B']],
                         height_ratios=[1, 1.4],
                         figsize=FIGURE_SIZES['spec_curve'])
ax_curve, ax_matrix = axes['A'], axes['B']
focal, ref = get_emphasis_pair()

# Sort specs by effect
order = np.argsort(effects)

# Spec curve (top panel)
for i, idx in enumerate(order):
    color = focal if sig[idx] else ref
    ax_curve.errorbar(i, effects[idx],
                      yerr=[[effects[idx] - lo[idx]],
                            [hi[idx] - effects[idx]]],
                      fmt='o', color=color, markersize=2,
                      elinewidth=0.5)
add_reference_line(ax_curve, value=0, orientation='h')
ax_curve.set_xlabel('')
ax_curve.set_ylabel('Effect size')

# Descriptor matrix (bottom panel)
for j, choice in enumerate(choices):
    for i, idx in enumerate(order):
        if choice in spec_choices[idx]:
            ax_matrix.scatter(i, j, marker='|',
                              color=focal if sig[idx] else ref,
                              s=20)
ax_matrix.set_yticks(range(len(choices)))
ax_matrix.set_yticklabels(choices)
ax_matrix.set_xlabel('Specification (ordered by effect)')

APA notes:

  • Report total n_specs, n_sig, n_null in the caption
  • State which decisions were preregistered vs exploratory
  • For a Bayesian variant, use a BF threshold (e.g., BF₁₀ > 3) instead of p-value cutoffs

Common mistakes:

  • Hiding any "garden of forking paths" decision (defeats the purpose)
  • Failing to label what specifications differ on (the matrix is the payload — without it, the curve is just a histogram)
  • Too many specs without ordering (visual mess; sort by effect)
  • Forgetting the null reference line

Representative paper: Botvinik-Nezer et al. (2020) NARPS, Nature; Simonsohn, Simmons & Nelson (2020), Nature Human Behaviour; Steegen et al. (2016), Perspectives on Psychological Science.

Worked example: see examples/figure_16_specification_curve.py.


33. Causal DAG

When to use: Display the assumed causal structure for a research question — treatments, outcomes, confounders, mediators, colliders. Standard in epidemiology (M-bias avoidance), economics (IV identification), and methodology (adjustment set choice). Anchored in Pearl (2009), Greenland, Pearl & Robins (1999), and the dagitty literature.

DAGs in matplotlib use pure FancyBboxPatch + FancyArrowPatch constructs — no dagitty / pgmpy / dowhy required for static publication-quality drawing. Those packages can compute d-separation and adjustment sets, but the static figure is plain matplotlib.

Key parameters:

  • Size: FIGURE_SIZES['dag']
  • Palette: NEUTRAL only — DAGs are typically monochrome; emphasis is via line weight or node shape (Ellipse for latent, FancyBboxPatch for observed), not color
  • Nodes: white fill, black edge; observed = rounded rectangle, latent / unobserved = ellipse
  • Arrows: FancyArrowPatch arrowstyle='->'; solid lw=1.6 for assumed causal arrows; dashed lw=1.0 for hypothesized / contested
  • Outcome node: thicker edge (lw=1.5) or a light fill to distinguish

Core code pattern:

from matplotlib.patches import FancyBboxPatch, FancyArrowPatch, Ellipse

fig, ax = plt.subplots(figsize=FIGURE_SIZES['dag'])
ax.set_xlim(0, 1); ax.set_ylim(0, 1); ax.axis('off')

# Nodes — (x, y, label, kind) for each
nodes = {
    'A': (0.15, 0.5, 'Treatment', 'box'),
    'Y': (0.85, 0.5, 'Outcome', 'box'),
    'L': (0.5, 0.85, 'Confounder', 'box'),
    'U': (0.5, 0.15, 'Unmeasured', 'ellipse'),  # latent
}
for name, (nx, ny, label, kind) in nodes.items():
    if kind == 'ellipse':
        ax.add_patch(Ellipse((nx, ny), width=0.18, height=0.10,
                             facecolor='white', edgecolor='black',
                             linewidth=1.0))
    else:
        ax.add_patch(FancyBboxPatch((nx - 0.09, ny - 0.05),
                                     0.18, 0.10,
                                     boxstyle='round,pad=0.02',
                                     facecolor='white',
                                     edgecolor='black', linewidth=1.0))
    ax.text(nx, ny, label, ha='center', va='center', fontsize=9)

# Directed edges — (from, to, 'solid'/'dashed')
edges = [
    ('A', 'Y', 'solid'),   # assumed causal effect
    ('L', 'A', 'solid'),
    ('L', 'Y', 'solid'),
    ('U', 'A', 'dashed'),  # hypothesized confounding
    ('U', 'Y', 'dashed'),
]
for start, end, style in edges:
    p_start = nodes[start][:2]
    p_end = nodes[end][:2]
    ax.add_patch(FancyArrowPatch(p_start, p_end,
                                 arrowstyle='->', mutation_scale=14,
                                 linewidth=1.4 if style == 'solid'
                                                else 1.0,
                                 linestyle='-' if style == 'solid'
                                                else '--',
                                 color='black'))

APA notes:

  • Latin variable names italic (per APA convention); Greek not italic
  • State which arrows are theoretical and which are empirically supported
  • If the figure depicts an adjustment set, indicate it explicitly (e.g., shaded boxes around the variables that must be controlled)
  • Reference Pearl (2009) and Greenland-Pearl-Robins (1999) when the DAG anchors a causal identification argument

Common mistakes:

  • Bidirectional arrows for "association" — this violates DAG convention. Only directed (single-headed) arrows belong in a DAG; covariation goes via a shared parent
  • Cycles — a DAG must be acyclic by definition. If the design suggests reciprocal causation, switch to a structural model (§9) with explicit time indices, not a DAG
  • Missing unmeasured confounders (U) when they are part of the identification argument
  • Confusing path coefficient diagrams (§9 path / SEM) with causal DAGs (§33) — they are different formal objects. A path diagram is a model of a fitted statistical structure; a DAG is a model of the assumed data-generating process

Representative paper: Pearl (2009) Causality; Greenland, Pearl & Robins (1999), Epidemiology; dagitty.net reference DAGs across epidemiology and economics.


Cross-references

  • For core distribution and bivariate charts (§1-§7), see chart-types-core.md.
  • For causal-econ designs (DID, event-study, RD, binscatter, multi-model coefficient panels), see chart-types-causal-econ.md. The IV diagram (§38) lives in §9.1 above and is referenced from there.
  • For raincloud, funnel, Kaplan-Meier, Bland-Altman, ridgeline, choropleth, ROC + calibration, mobility matrix, and marginal-effects prediction, see chart-types-applied.md.
  • For the 11 P3 speed-look entries (histogram, KDE, swarm, factor loading, Johnson-Neyman, CONSORT, lollipop, slope, dumbbell, pyramid, tornado), see chart-types-quick-reference.md.
  • For panel composition, shared axes, and multi-panel labeling (formerly §13 in this guide), see multi-panel.md.
  • For palette selection, see color-system.md; for APA conventions including §11 Mediation Diagram Conventions, see apa-figure-standards.md.