News · Brain & Mental Health
Practice at brain tests hides a year of cognitive decline
Older adults get better at cognitive tests simply by taking them again. In 977 people across 2,571 assessments, that practice gain was large enough to cancel out about a year of age-related decline.
- Cognitive decline is measured by testing the same people repeatedly over years.
- People improve at those tests from familiarity alone, which is called a retest effect.
- Across 977 older adults, the gain offset roughly one year of age-related decline.
- It showed up in global cognition, executive function and memory, but not in language.
- Adjusting for it changed the decline curves and did not change the trial's conclusion.
Almost everything known about how fast the aging brain declines comes from testing the same people over and over. Word lists, number sequences, naming tasks, repeated every year or two for a decade.
There is a problem built into that method, and it has been known for years without being routinely handled. People get better at tests they have taken before.
An analysis in JAMA (Journal of the American Medical Association) Network Open measured how much better, in 977 older adults across 2,571 assessments. The practice gain was roughly the size of a year of decline.
What cognitive tests are meant to do
Mental status testing is done to check a person’s thinking ability and to determine if any problems are getting better or worse.
It is not one test but a battery covering appearance, attitude, orientation, psychomotor activity, attention span, memory both recent and long-term, language function, and judgment and intelligence. Most tests are divided into sections, each with its own score, and the results help show which part of someone’s thinking and memory may be affected.
The MedlinePlus guidance carries a caution worth keeping in view: an abnormal mental status test alone does not diagnose the cause.
This study adds a second caution, about the opposite result. A test that has not got worse may not mean what it appears to.
Why repeated testing flatters people
Retest effects are improvements in test performance from repeated exposure, potentially masking true cognitive decline by offsetting approximately one year of age-related cognitive loss.
The mechanism is not mysterious. On a second sitting you recognize the format, you remember some of the items, and you have a strategy for the task that you had to invent the first time. None of that reflects a change in the brain being measured.
The question this analysis asks is not whether that happens. It is how big the effect is, whether it differs across the things being measured, and whether it distorts the results of trials that use cognitive decline as an outcome.
How the retest effect was measured
The data come from an ad hoc secondary analysis of ACHIEVE, a multicenter phase 3 randomized clinical trial conducted from 2017 to 2022 with 3-year follow-up. The ACHIEVE study evaluated whether a hearing intervention could mitigate cognitive decline compared with health education control in older adults with untreated hearing loss.
A total of 977 older adults (mean baseline age, 76.3 years; 523 [54%] female) contributed 2571 in-person cognitive assessments across four US community sites.
The trick is in how retesting was modeled. Cognitive retesting was defined as a binary indicator (0 for baseline visit; 1 for follow-up visits), which separates the one-off jump that comes from having done the thing before from the gradual change that comes with age.
The primary outcome was a global cognitive factor score derived from a neurocognitive battery, with secondary outcomes covering language, executive function and memory separately.
How big the practice gain was
Retest differences were significant for global cognition, executive function, and memory, and the scale of that is the finding.
Those gains were approximately offsetting 1 year of model-estimated age-related cognitive decline. In a study measuring change over three years, one year of it can be practice.
The exception is informative. Language showed minimal retest differences, with an estimate essentially at zero.
That split makes sense of the pattern. Executive function and memory tasks reward strategy and familiarity; naming and vocabulary tasks draw on knowledge a person either has or does not. The tests most used to detect early decline are the ones most inflated by having sat them before.
The effect was also stubbornly uniform across people. Retest differences did not meaningfully vary by hearing loss severity, trial intervention, or recruitment source.
Why retest effects did not overturn the trial
The reassuring half of the paper is what did not change. Retest adjustment altered estimated cognitive trajectories but did not meaningfully alter intervention outcomes.
The reason is structural. In a randomized trial, both arms take the same tests the same number of times, so the practice gain lands equally on both and largely cancels when you compare them.
Where it does not cancel is in describing a single group. A study reporting that a cohort declined slowly, or that people doing some activity held steady, has that same inflation baked in with nothing to subtract it against.
The authors’ recommendation is procedural rather than dramatic: adjustment for retest differences is recommended to improve interpretation and intervention estimation in longitudinal cognitive studies.
What this retest analysis cannot show
This was not planned in advance. An ad hoc secondary analysis asks a question of data collected for something else, which is the right way to investigate a measurement problem and a weaker footing than a designed study.
There is also a design gap the authors record: year 2 assessments were excluded by design because most were telephone-based during the COVID-19 period and were not directly comparable.
And the participants are specific: older adults recruited for a hearing trial, mean age 76, at four US sites. Whether the same practice gain applies at 60, or in a memory clinic, is not established here.
What to ask of the next brain study you read
The practical use of this is as a question to carry into other coverage.
When a study reports that some group declined more slowly than expected, or that an activity appeared to protect cognition, it is worth knowing how many times those people sat the test and whether the analysis accounted for it. On these numbers, a year of apparent preservation can be an artifact of familiarity.
The same applies at an individual level. If a relative is having repeat cognitive assessments, a score that holds steady is a better sign than a score that falls, and it is not the same as a brain that has not changed.
The same trap catches supplement trials. A first-in-human test of strawberry leaf extract reported within-group memory gains across three sittings of the same battery, which is exactly the shape practice produces.
People also ask
What is a retest effect?
Retest effects are improvements in test performance from repeated exposure, potentially masking true cognitive decline by offsetting approximately one year of age-related cognitive loss. In plain terms: you get better at a test because you have done it before, not because your brain has improved.
How large were the effects here?
Retest differences were significant for global cognition (beta 0.11; 95% CI 0.07 to 0.15), executive function (0.15; 0.10 to 0.20), and memory (0.12; 0.06 to 0.18), approximately offsetting 1 year of model-estimated age-related cognitive decline. Language showed minimal retest differences (-0.03; -0.08 to 0.02).
Why does language behave differently?
The study does not test why. A plausible reading is that vocabulary and naming tasks depend on knowledge a person already has rather than on a strategy that can be learned across sittings, so there is less room to improve by familiarity. That is an interpretation, not a finding.
Did it change the trial's answer?
No. Retest adjustment altered estimated cognitive trajectories but did not meaningfully alter intervention outcomes. That is the reassuring half: because both arms of a randomized trial get the same practice, the effect largely cancels when you compare them. It does not cancel when you are describing how fast a group declined.
What was the ACHIEVE trial?
The ACHIEVE study evaluated whether a hearing intervention could mitigate cognitive decline compared with health education control in older adults with untreated hearing loss. This analysis is an ad hoc secondary look at its cognitive data, not a test of hearing aids.
Who was in it?
977 older adults with a mean baseline age of 76.3 years, 523 (54%) female, contributing 2,571 in-person cognitive assessments across four US community sites. 552 (57%) had mild hearing loss and 739 (76%) were recruited de novo rather than from an existing cohort.
What should a reader do with this?
Treat it as a lens for reading other studies rather than as advice. This is general information rather than medical advice. When a study reports that a group declined slowly, or that an activity slowed decline, the question worth asking is whether repeated testing was accounted for. And if you or a relative are having repeat cognitive assessments, a stable score is not necessarily a stable brain.