What Is the Testing Effect? Six Tests on the Same 5,021 Cards
The testing effect means one retrieval outlasts one re-reading. We tracked the same 5,021 vocabulary cards across six tests: the gaps kept growing, yet scores rose from 11.8% to 77.9%.
What Is the Testing Effect? Six Tests on the Same 5,021 Cards
Testing is usually treated as a way to check whether you have learned something. But cognitive psychology has a result that keeps replicating: the test itself is one of the most effective learning actions available, more effective than reading the material again. This is the testing effect.
We tracked the first six tests on the same 5,021 vocabulary cards. Every gap was longer than the one before, so each test was harder than the last — and yet the scores rose each time, from 11.8% to 77.9%.
What the testing effect is
The testing effect (also called the retrieval practice effect) is this: retrieving a piece of information from memory once does more for long-term retention than re-reading the same information once.
The most cited evidence is the 2006 experiment by Roediger and Karpicke. Participants read a passage and were then split into two groups: one re-read it, the other took a recall test. On an immediate test the re-reading group did better; a week later, the tested group was clearly ahead.
Two conclusions follow:
- A test is not just a measurement. It changes the memory itself.
- How it feels in the moment misleads you. Re-reading feels smoother and performs worse in the long run.
How it relates to active recall
The two terms get used interchangeably, but they sit at different levels:
| Active recall | Testing effect | |
|---|---|---|
| What it is | A study action | An experimental result |
| Content | Retrieve first, then check | Retrieval outlasts re-reading |
| Question it answers | What should I do | Why should I do it that way |
Active recall is the method; the testing effect is why the method works.
The data: six tests on the same cards
The hard part about observing the testing effect is that difficulty swamps everything — cards that get tested many times are the difficult cards to begin with. Pool every card together and what you see is a difference in difficulty, not a learning curve.
So we fixed the population: we took only cards tested exactly 6 to 8 times and followed their first six tests. Every row in the table below describes the same set of cards, not different ones.
| Item | Value |
|---|---|
| Period | 13 April – 23 August 2026 |
| Cards | 5,021 |
| Observations per row | 7,724 |
| Test number | Gap since last | Recalled clearly |
|---|---|---|
| 1st | same day | 11.8% |
| 2nd | 0.13 days | 49.6% |
| 3rd | 0.51 days | 59.3% |
| 4th | 1.27 days | 64.9% |
| 5th | 2.27 days | 71.5% |
| 6th | 5.25 days | 77.9% |
The middle column is the point. From the 1st test to the 6th, the gap since the previous test grew from same-day to 5.25 days. The longer the gap the harder retrieval becomes (see the forgetting curve), so every test was harder than the one before it.
With the questions getting steadily harder, the score climbed from 11.8% to 77.9%.
It holds for hard cards too — it is just slower
Splitting the cards into four groups by how many times they were tested in total — more tests means the card was harder for that user — every group shows the same upward trend:
| Card difficulty | 1st test | Midpoint | Final tracked test |
|---|---|---|---|
| Tested 3–5 times (easiest) | 39.7% | 75.8% (2nd) | 83.3% (3rd) |
| Tested 6–8 times | 11.8% | 64.9% (4th) | 77.9% (6th) |
| Tested 9–12 times | 6.3% | 52.8% (5th) | 73.4% (9th) |
| Tested 13+ times (hardest) | 4.7% | 39.5% (6th) | 62.8% (13th) |
The hardest group came back only 4.7% of the time on the first test — effectively a complete blank. By the 13th test it reached 62.8%.
The slope flattens as difficulty rises, but the direction has no exceptions.
One caveat worth stating: every encounter in this data is "try first, then check the answer" — there is no control group that only re-read without testing. So what we measured is the cumulative effect of repeated retrieval, not a direct comparison between retrieval and re-reading. That comparison requires randomised laboratory work, which is exactly what the Roediger and Karpicke line of research provides.
Three things you can act on
1. Blanking on the first attempt is the default, not bad news. The hardest group scored 4.7% on the first test. Holding yourself to "I should get it right the first time" makes people quit before they have started.
2. Put the test right after the reading, not at the very end. The common approach is to read the whole book and then work through questions to check. The reverse works better: read a small section, close it, and quiz yourself immediately. Every retrieval consolidates, so the earlier you start, the more consolidations you accumulate.
3. Fluency is not a reliable signal. Re-reading feels competent and testing feels laborious, while the long-term results run the other way. That is exactly why so many people know they should self-test and go back to re-reading anyway.
FAQ
What is the testing effect?
The testing effect is the finding that retrieving information from memory once helps long-term retention more than re-reading the same information once. It was established by the 2006 experiment of Roediger and Karpicke: the group that took a recall test after reading clearly outperformed the group that re-read a week later, even though the re-reading group did better on an immediate test.
How is the testing effect different from active recall?
Active recall is a study action — retrieve the answer yourself, then check it. The testing effect is the experimental result explaining why that action works: retrieval outlasts re-reading. One is the method, the other is the reason the method works.
Why do tests help memory?
Because the act of retrieval changes the memory and makes the next retrieval easier, rather than simply reading out what is already stored. On top of that, a test exposes on the spot which items you do not actually know, so the time that follows goes into the real gaps.
Is there data supporting the testing effect?
Yes. Alongside laboratory research, we tracked the first six tests on the same 5,021 vocabulary cards: the share recalled clearly rose from 11.8% on the 1st test to 77.9% on the 6th, while the gap between tests grew from same-day to 5.25 days. The questions got harder and the scores got better.
Is it normal to draw a complete blank the first time?
Yes. We split cards into four difficulty groups: the hardest group scored only 4.7% on the first test, and even the easiest group scored 39.7%. All four groups then rose steadily, with the hardest reaching 62.8% by its 13th test. First-attempt performance does not predict the final outcome.
When should I test myself?
Right after finishing a small section, rather than saving it all up until you have read everything. Every retrieval consolidates the memory, so starting earlier means more retrievals fit into the same amount of time.
Does the testing effect apply to vocabulary?
It does, and vocabulary is the most direct application, because every word is a separately retrievable item. A flashcard is the testing effect turned into a procedure: the question appears on the front and the answer stays hidden until you have produced it.
Why does re-reading feel more effective?
Because re-reading is smooth and testing is effortful, and people mistake smoothness for having learned something. That mismatch is the most counter-intuitive part of the testing effect: the method that feels worse in the moment produces the better long-term result.
Should the testing effect be combined with spaced repetition?
Yes. The testing effect decides what you do during a review — retrieve, rather than re-read — and spaced repetition decides when that action happens. Together they compound: a single retrieval placed just before you would forget costs the least and consolidates the most.
The data in this article comes from anonymised review records on Tomaru. Further reading: What is active recall, What is the forgetting curve.