If you take the 1998 NAEP 8th grade reading scores by race, re-weight them by the new 2024 demographic makeup of American 8th graders, you will find a predicted point decline of 4.6, rather close to the actual 4 point fall over the period.
Over the same period, reading scores actually improved somewhat for each of the 1998 poorly scoring groups (and quite a bit for Asians), but the poorly scoring groups simply make up so much more of the population now that it drags everything else down.
1998? I wanted a quarter century of Grade 8 and that gives me 1998 to 2024 with NAEP. 2000 reading assessment was Grade 4 only. You can do it with any pair of years.
The article says that scores began drifting down around 2012. Gemini is telling me that NAEP peaked in 2013. Elsewhere in this thread there's a claim of an inflection point around 2018 in PISA, which eyeballing Gemini's numbers lines up with NAEP.
Does this correlate with the racial proportion data?
If you graph actual NAEP since 1998 versus a 1998 group-fixed score, you'll see increases (1998-2002, 2005-2013, 2015-2017) and multiple declines twice taking the line back down to the 1998 linear trend prediction (2002-2005, 2017-2024).
It is wrong to get overly exact here though. We are talking about a overall score variance of just 10pts in the face of an achievement gap of nearly 30pts, with an absolutely massive 20% decline in white population share over the 26 year period.
If you're seeing trends in your residuals that's a violation of the independence assumption of linear models (assuming you're using about a linear model since you're talking about "linear predictions").
I'm in a different field but an arbitrary 27 years (1998-2024) of data and the autocorrelation both would get flagged in a review for me. Not to give statistics homework but you should test and correct the autocorrelation issue you're describing in the model and if the data goes farther back I would go farther back, too (with how you're describing the errors going off pattern then back on it sounds like the inference here would be at least somewhat unstable depending on year chosen).
Edit: these are the assumptions and basics of how to correct for violations, in particular you're describing a violation of assumption 2 but you should test for all of them - <a href="https://www.statology.org/linear-regression-assumptions/" rel="nofollow">https://www.statology.org/linear-regression-assumptions/
I was checking if there was a response and upon reading your comments again, looks like your method is less a model and more just some ad hoc procedure. Based on what your conclusion is it sounds like a linear model based on demographic data would work and linear models are easy to train. Though by your summary there are probably going to be issues with the model and you have the same issues with your procedure, you just can't tell because it hasn't been rigorously specified.
There's a channel on YouTube called Statquest that teaches statistics in a pretty accessible way if you're interested in analyzing this sort of data.
I learned my statistics in university, not Youtube. You are confusing yourself about residuals.
To make my point in the original post, I held the group scores fixed and extrapolated combined scores based on changing population sizes. This is straightforward. Simple even.
\(\widehat S=\sum_g p_gS_g,\)
Why did I call this "linear"? (Not a "linear model", your term, because it is not). The model function actually tracks the population function. Over the 25-year period, however, this trend is in fact linear because the White-to-Hispanic population trend is linear on that timescale. Over a longer timescale it would not be.
What you're describing is a linear model, it just has an ad hoc construction instead of something like a least squares training (<a href="https://en.wikipedia.org/wiki/Linear_model" rel="nofollow">https://en.wikipedia.org/wiki/Linear_model). Also linear models are about linear associations and work even with non-linear predictors. Your assumption in the model is explicitly a linear association so how the predictors vary shouldn't affect it.
I don't mean to disparage amateur statistics since often there are interesting things that experts miss that amateurs have insight into but this sort of data with few data points (the data you have doesn't sound independent) and a lot of confounding factors is not trivial to accurately model. Not that everyone needs to learn statistics but this sort of data (and also this topic) probably deserves a bit more of a rigorous approach.
As a very simple example, picking a time range for the model because the model doesn't work when you go farther back adds a lot of potential bias into the model and that along with the non-independence issue (and without looking at the data myself I don't know if there are other issues) are going to lead to overconfidence in conclusions.
romaaeterna · · focus · HN ↗
Over the same period, reading scores actually improved somewhat for each of the 1998 poorly scoring groups (and quite a bit for Asians), but the poorly scoring groups simply make up so much more of the population now that it drags everything else down.
| Group | 1998 NAEP | 2024 NAEP | Change |
|---|---:|---:|---:|
| White | *268* | *266* | −2 |
| Black | *242* | *243* | +1 |
| Hispanic | *241* | *245* | +4 |
| Asian/Pacific Islander | *261* | *280* | +19 |
| All public-school students | *261* | *257* | −4 |
| Race/ethnicity | 1999 Grade 8 | 2023 Grade 8 | Change |
|---|---:|---:|---:|
| White | *64.5%* | *44.0%* | −20.5 pp |
| Black | 16.1% | 14.9% | −1.2 pp |
| Hispanic | *14.3%* | *29.4%* | +15.1 pp |
| Asian/Pacific Islander* | 3.9% | 5.9% | +2.0 pp |
| American Indian/Alaska Native | 1.2% | 0.9% | −0.3 pp |
| Two or more races | not separately reported | 4.8% | — |
tdb7893 · · focus · HN ↗
romaaeterna · · focus · HN ↗
1998? I wanted a quarter century of Grade 8 and that gives me 1998 to 2024 with NAEP. 2000 reading assessment was Grade 4 only. You can do it with any pair of years.
kalkin · · focus · HN ↗
Does this correlate with the racial proportion data?
romaaeterna · · focus · HN ↗
It is wrong to get overly exact here though. We are talking about a overall score variance of just 10pts in the face of an achievement gap of nearly 30pts, with an absolutely massive 20% decline in white population share over the 26 year period.
tdb7893 · · focus · HN ↗
I'm in a different field but an arbitrary 27 years (1998-2024) of data and the autocorrelation both would get flagged in a review for me. Not to give statistics homework but you should test and correct the autocorrelation issue you're describing in the model and if the data goes farther back I would go farther back, too (with how you're describing the errors going off pattern then back on it sounds like the inference here would be at least somewhat unstable depending on year chosen).
Edit: these are the assumptions and basics of how to correct for violations, in particular you're describing a violation of assumption 2 but you should test for all of them - <a href="https://www.statology.org/linear-regression-assumptions/" rel="nofollow">https://www.statology.org/linear-regression-assumptions/
tdb7893 · · focus · HN ↗
There's a channel on YouTube called Statquest that teaches statistics in a pretty accessible way if you're interested in analyzing this sort of data.
romaaeterna · · focus · HN ↗
To make my point in the original post, I held the group scores fixed and extrapolated combined scores based on changing population sizes. This is straightforward. Simple even.
\(\widehat S=\sum_g p_gS_g,\)
Why did I call this "linear"? (Not a "linear model", your term, because it is not). The model function actually tracks the population function. Over the 25-year period, however, this trend is in fact linear because the White-to-Hispanic population trend is linear on that timescale. Over a longer timescale it would not be.
tdb7893 · · focus · HN ↗
I don't mean to disparage amateur statistics since often there are interesting things that experts miss that amateurs have insight into but this sort of data with few data points (the data you have doesn't sound independent) and a lot of confounding factors is not trivial to accurately model. Not that everyone needs to learn statistics but this sort of data (and also this topic) probably deserves a bit more of a rigorous approach.
As a very simple example, picking a time range for the model because the model doesn't work when you go farther back adds a lot of potential bias into the model and that along with the non-independence issue (and without looking at the data myself I don't know if there are other issues) are going to lead to overconfidence in conclusions.
romaaeterna · · focus · HN ↗