QML - Week 6

Regression models: the basics

Stefano Coretta

Word frequency and reaction times

What is the relationship between a word’s lexical frequency and reaction times in a lexical decision task in Croatian?

Which relationship?

Reaction times

Word frequency and RTs

Logged word frequency and RTs

Gaussian model of RT

\[RT \sim Gaussian(\mu, \sigma)\]

But we want to know what happens to RTs depending on the value of lexical frequency…

Then we let the mean \(\mu\) vary by lexical frequency!

\[\begin{aligned} RT_i & \sim Gaussian(\mu_i, \sigma)\\ \mu_i & = \beta_0 + \beta_1 \cdot logf_i \end{aligned}\]

The equation of a line

\[y = \beta_0 + \beta_1 \cdot x\]

  • \(\beta_0\) is the line intercept: the \(y\) value when \(x\) is 0 zero.

  • \(\beta_1\) is the line slope: the change in \(y\) for each unit-increase of \(x\).

Regression model

\[\begin{aligned} RT_i & \sim Gaussian(\mu_i, \sigma)\\ \mu_i & = \beta_0 + \beta_1 \cdot logf_i & \text{[Regression equation]} \end{aligned}\]

  • A regression model is a model that uses the equation of a line (the regression equation).

  • The model estimates \(\beta_0\) (the intercept), \(\beta_1\) (the slope) and \(\sigma\) from the data (i.e. the observed \(RT\) and \(logf\) values).

\(\beta_0\), intercept

  • Mean RT value when logged frequency is 0 zero (i.e. when word frequency is 1; exp(0) = 1).

\(\beta_1\), slope

  • Change in mean RT for each unit increase of log-frequency (when log-frequency goes from \(x\) to \(x + 1\)).

Fit model in brms

rt_bm <- brm(
  rt ~ 1 + log_freq,
  family = gaussian,
  data = croat,
  cores = 4,
  seed = 4125,
  file = "cache/rt_bm"
)

Model results

library(bayesplot)

mcmc_dens(rt_bm, pars = c("b_Intercept", "b_log_freq", "sigma"))

Posterior draws

rt_draws <- as_draws_df(rt_bm)
rt_draws
# A draws_df: 1000 iterations, 4 chains, and 6 variables
   b_Intercept b_log_freq sigma Intercept lprior   lp__
1         1096        -44   107       664    -11 -15847
2         1120        -47   103       667    -11 -15846
3         1095        -45   105       660    -11 -15846
4         1121        -47   104       667    -11 -15845
5         1103        -45   102       663    -11 -15845
6         1105        -45   106       665    -11 -15845
7         1110        -46   103       664    -11 -15844
8         1092        -44   103       663    -11 -15845
9         1125        -47   103       666    -11 -15846
10        1114        -46   104       662    -11 -15845
# ... with 3990 more draws
# ... hidden reserved variables {'.chain', '.iteration', '.draw'}

Posterior of \(\beta_1\)

rt_draws |> 
  ggplot(aes(b_log_freq)) +
  geom_density(fill = "brown", alpha = 0.5)

Posterior of \(\beta_1\) with CrIs

CrIs of \(\beta_1\)

library(posterior)

# 95%
quantile2(rt_draws$b_log_freq, c(0.025, 0.975)) |> round()
 q2.5 q97.5 
  -48   -43 
# 80%
quantile2(rt_draws$b_log_freq, c(0.1, 0.9)) |> round()
q10 q90 
-47 -44 
# 60%
quantile2(rt_draws$b_log_freq, c(0.2, 0.8)) |> round()
q20 q80 
-47 -45 

Expected values

\[\begin{aligned} \mu_i & = \beta_0 + \beta_1 \cdot logf_i\\ \mu_{[logf = 4]} & = \beta_0 + \beta_1 \cdot 4\\ \mu_{[logf = 12]} & = \beta_0 + \beta_1 \cdot 12\\ \end{aligned}\]

rt_draws <- rt_draws |> 
  mutate(
    rt_logf_4 = b_Intercept + b_log_freq * 4,
    rt_logf_12 = b_Intercept + b_log_freq * 12
  )

quantile2(rt_draws$rt_logf_4, c(0.1, 0.9)) |> round()
q10 q90 
916 933 
quantile2(rt_draws$rt_logf_12, c(0.1, 0.9)) |> round()
q10 q90 
556 564 

Plotting expected values

Word frequency and reaction times (bis)

What is the relationship between a word’s lexical frequency and reaction times in a lexical decision task in Croatian?

  • When log-frequency is 0, the mean RTs are between 1084 and 1129 ms at 95% confidence.

  • For each unit increase of log-frequency, the mean RTs decrease by 43-48 ms, at 95% confidence.

Correlation in NOT causation

Be careful!

  • Correlation between two variables: they co-vary, i.e. they show a systematic association (their values tend to vary together in a consistent pattern).

  • Spurious correlations: two variables can be correlated because of bias from another variable.

  1. Number of plant names in a language vs. biodiversity of the region
    • Languages in biodiverse regions have more words for plants.
    • Mediator: cultural reliance on plants.
  2. Language endangerment vs. economic development
    • Higher economic development is associated with greater language endangerment.
    • Confounder: colonial history.
  3. Language prestige vs government policy
    • High prestige languages and officially supported languages each attract learners.
    • Collider: If you only look at languages with many learners, prestige and policy might appear related even if they’re not causally connected.

But it is if you use causal inference…

Causal inference

  • Correlation can be interpreted causally if you adopt a causal inference approach.

  • Learn about it in McElreath’s textbook Statistical Rethinking. Also check STeW.