You reach the front of a flashcard and the answer arrives before you have finished reading. Good sign. Probably.
Then a lecturer asks about the same idea in a different sentence. Or an exam gives you a new case instead of the definition you memorized. The answer, so cooperative a moment ago, becomes strangely unavailable.
A card can be easy for two reasons. It can be easy because you know the answer. It can also be easy because you know the card.
The title is a provocation, not a statistic: the research reviewed here does not estimate how many real decks are badly designed. It shows something narrower: success under one familiar cue may not survive a new prompt or task.
To write a good flashcard, give memory one clear job: put a specific prompt on the front, the smallest complete and checkable answer on the back, and enough context to identify the question without letting the wording answer it for you. Split a card when its parts can fail independently. Change the practice entirely when the real goal is extended reasoning, problem-solving, synthesis, or performance.
Recognition can teach. It still proves less
Recognition means identifying an answer already in view. Cued recall provides a prompt but requires you to produce the target. Free recall asks more broadly what you remember.
A normal flashcard is a cued-recall task, so it needs a cue. The design question is how to specify the job without quietly doing the job.
In a classic experiment by Henry Roediger and Jeffrey Karpicke, students either restudied prose passages or repeatedly recalled them without feedback. Five minutes later, the restudy group remembered 81 percent and the repeatedly tested group 75 percent. One week later, the order had reversed: repeated testing produced 61 percent recall, compared with 40 percent after repeated study. Smooth performance during practice and durable learning had come apart.
Recognition can also give a correct answer when production fails. In one unusual multiple-choice experiment, students sometimes correctly selected "none of the above" and were then asked to supply the answer. In 45 percent of those cases, they wrote the wrong answer or nothing at all. This was not a flashcard study, and it does not mean that 45 percent of card successes are false. It demonstrates a narrower point: identifying the right response can coexist with being unable to generate it.
Multiple-choice practice is not worthless. Across four experiments with 372 learners, Megan Smith and Jeffrey Karpicke found that every retrieval format beat study alone after a week. Short-answer and hybrid practice showed little or no advantage in three experiments; an advantage appeared only when initial recall success improved. Harder was not automatically better.
A recognition question can strengthen memory. It simply gives weaker evidence that the learner could have produced the answer without the options, wording, or familiar setting.
The broader case for retrieval, spacing, and feedback is explained in why flashcards work. Here, the narrower question is whether the card itself asks for the knowledge or merely points at it.
How to write flashcards that give memory one clear job
Consider this card:
Explain photosynthesis.
It may be a reasonable essay prompt. As a routine flashcard, it is difficult to score. Does a correct answer require the overall equation, the light-dependent reactions, the Calvin cycle, or the origin of the oxygen? You can remember half the process and still have no principled way to choose between "Knew it" and "Forgot."
Broad retrieval is not inherently bad. In a study of 103 people reading an instructional text about coffee, specific questions helped after a week when the text contained interesting but irrelevant details. Without those details, broad and specific prompts performed similarly. Specificity helped direct retrieval toward the intended material.
That is why "one card, one target" is best treated as a design rule, not a law discovered by cognitive science. A target can be a fact, a distinction, a short mechanism, or a decision. It need not fit into one word.
A practical test is independence of failure. If two parts of an answer can be forgotten separately and need different correction, they probably deserve separate cards. Define reliability, validity, bias, and confounding hides four possible gaps inside one self-rating. By contrast, Why does increasing sample size usually reduce standard error? may need several sentences, but they belong to one causal explanation.
A flashcard answer should be exactly as long as it takes to complete one clearly defined target — and short enough that you can judge it consistently. The point is diagnostic clarity, not shortness.
Give the question enough context, then stop
Put one identifiable retrieval target on the front. On the back, include the smallest complete answer needed to judge the attempt, together with any condition or exception that changes its meaning.
That sounds simple. It becomes less simple when the card either says too much or almost nothing.
1948 → ? may refer to the Universal Declaration of Human Rights, the Berlin Blockade, the founding of the NHS, or something meaningful only because of the previous card. The learner must first guess which question the author intended.
A useful cue removes uncertainty about the target while keeping the answer absent:
What document did the UN General Assembly adopt in 1948 as a statement of fundamental human rights?
The opposite problem is answer leakage. A question may echo the textbook so closely that its first words trigger the rest. Grammar, a half-label, color coding, or fixed card order can also supply accidental clues.
Images deserve the same suspicion. They help when the image itself is the knowledge target: an anatomical structure, a graph, a molecule, or an unlabeled diagram. Decorative pictures do not make a card more active. Labels, arrows, color conventions, and repeated layouts can also become answer leakage with better production values.
Asher Koriat and Robert Bjork found that judgments of future memory become optimistic when people assess learning while information is present that will be absent at test. They called this a foresight bias. Retrieval fluency can mislead for a related reason: an answer that comes quickly feels secure, although ease is not an infallible measure of later access.
Still, cues are not the enemy. In a study of arbitrary image-word pairs, progressively revealing the cue produced 39 percent recall after 48 hours, compared with 33 percent after standard retrieval and 26 percent after restudy. The task was unlike a normal text card, so the percentages should not be generalized directly. The result does warn against the belief that fewer cues must always be better.
A good cue tells you what to retrieve. A bad cue either leaves that unclear or supplies so much accidental information that little retrieval remains.
Cloze works when the blank is the idea
Are cloze flashcards effective? They can be, when the missing word, symbol, phrase, or formula component is itself the knowledge target. They provide much weaker evidence that you can explain or apply the idea surrounding the blank.
Cloze deletion is often accused of turning study into word guessing. Sometimes it does. Producing a missing item from a sentence can still be genuine recall. Trouble begins when deleting a word is mistaken for deciding what knowledge matters, or when the blank stands in for an idea the learner still can't explain or apply beyond it.
Scott Hinze and Jennifer Wiley tested fill-in-the-blank retrieval with prose. After delays of two and seven days, completion practice improved repeated questions but not related questions. In a third experiment, only broader paragraph recall improved transfer to novel questions. Narrow retrieval was effective at the narrow job it had been given.
This card may therefore be fine when the target is vocabulary:
The process by which plants convert light energy into chemical energy is called {{photosynthesis}}.
It tells us much less about whether the learner can explain the process. For that, ask another question:
What energy conversion takes place during photosynthesis?
Or test a mechanism:
What do ATP and NADPH from the light-dependent reactions provide to the Calvin cycle?
A middle-school science study found the same kind of asymmetry. Practice with application questions improved later definitions and new applications. Definition questions did not produce the same benefit on later application questions. Knowing what a principle means and knowing when to use it are related, but they are different retrieval jobs.
Five bad flashcard examples, repaired
1. The question is too broad
Bad: Explain photosynthesis.
Broad retrieval can be useful, but this card has no stable success criterion. Replace it with a defined relationship or mechanism:
What are the carbon source and energy source used to build glucose during photosynthesis?
Why can reduced light eventually limit glucose production even when carbon dioxide remains available?
Keep the broad prompt for blank-page recall or an oral explanation, where reconstruction is the point.
2. The answer contains independent ideas
Bad: Define reliability, validity, bias, and confounding.
Four gaps can hide inside one answer. Split it, and test boundaries or use:
How does reliability differ from validity?
What makes a third variable a confounder rather than merely a correlate?
In a voluntary online survey, what bias may arise from self-selection?
3. The wording does most of the work
Bad: The process by which plants convert light energy into chemical energy is called {{photosynthesis}}.
This is a reasonable vocabulary card. It is weak evidence that the learner understands photosynthesis. Keep it if the target is the term; add a mechanism or application card if the target is the process.
4. The cue is vague for no useful reason
Bad: 1948 → ?
Replace it with:
In what year was the Universal Declaration of Human Rights adopted?
Use the reverse direction only when the future task requires it:
What document did the UN General Assembly adopt in 1948 as a statement of fundamental human rights?
The two directions are different tasks. Automatically duplicating every card does not make a deck more complete.
5. The AI card says more than the source
Bad: What did the study prove? Social media causes depression.
Assume the source described an observational association. "Prove" erases uncertainty; "causes" strengthens the claim beyond the design; the population has disappeared.
Better cards would ask:
What association did the study report between social-media use and depressive symptoms in the sampled population?
Which feature of the observational design limits a causal interpretation?
What alternative explanation did the authors discuss?
This repair prevents the deck from teaching a stronger claim than the paper made.
Atomic cards can become atomized
"Atomic flashcards" usually means that each card has one independently assessable target. That helps self-grading and lets one weak idea return without dragging several remembered ideas along with it.
The studies cited here do not establish one fact per card, a maximum answer length, or a correct number of cards per chapter. Atomicity is a practical heuristic, not a law discovered by cognitive science. It fails when a deck is divided so finely that every relation between ideas disappears.
A deck can be wonderfully atomic and educationally pulverized.
Use narrow cards for prerequisite facts, then add distinctions, mechanisms, applications, and occasional larger reconstructions. Retrieval can transfer to new questions, but benefits often shrink when the final task changes format. In a classroom study of 182 eighth-graders, vocabulary retrieval helped across several later formats, yet the advantage was smaller when students had to reverse the cue, create a phrase, or retrieve from a new context.
So how many flashcards should you make? There is no useful universal number. Count worthwhile retrieval targets, not pages, and remember that every card creates future reviews. A card is a recurring appointment with your future self. Some appointments are not worth scheduling.
Some knowledge needs a different kind of practice
Flashcards are good at keeping components available: terms, formulas, criteria, distinctions, steps, and short causal links. They do not automatically assemble those components into the performance you eventually need.
If the exam requires worked problems, solve problems. If a viva requires a coherent explanation, practice explaining aloud. Use diagrams for spatial systems. Critique real study designs if you need to evaluate research. Rehearse a laboratory or clinical procedure in a form that includes the actual decisions and actions.
Practice components with cards, then practice the whole task in the form it will take. Research on completion questions, application questions, and changes in test format all points in that direction. Retrieval can make knowledge more portable, but portability depends partly on what was retrieved and what the later task demands.
AI-generated flashcards need an editor
Can AI create flashcards? Easily. It can also create a large deck before anyone has decided whether the questions deserve to exist. Speed moves the bottleneck from production to judgment.
A 2026 randomized study compared AI-assisted and student-written medical questions in mock examinations taken by 258 first-year students. The AI workflow was 5.6 times faster, averaging 4.2 minutes per item rather than 19.6. Students rated clarity, relevance, difficulty, and educational value similarly. Yet student-written items had slightly better discrimination, the exam was formative and open-resource, and the authors recommended a human-in-the-loop workflow.
Other source-grounded work finds that questions can be on topic and answerable while remaining shallow or needing expert filtering. A fluent question is not necessarily useful.
Reviewing an AI-generated deck starts with the target: what exactly must you produce, and can different parts of the answer fail independently? Then look for leakage in the wording, grammar, image, or card order. Check the answer against its source, including the population, conditions, and uncertainty of the original claim. Ask whether the question type matches the learning goal. Finally, inspect the deck as a whole: a card may be accurate and clearly written while still being duplicated, peripheral, or too trivial to deserve months of future reviews.
Quizpace can turn learner-provided material into draft flashcards and questions that learners review before relying on them.
Generation is the quick part. Judgment is still yours.
The five-second test
For any card, ask:
What exactly am I trying to recall?
Could I answer from the wording or format alone?
Can I judge the answer quickly and consistently?
Will I need this knowledge in this form later?
Would another kind of practice test it better?
A fast answer can mean that memory is strong. That is what practice should produce. But once a card becomes very familiar, change something that should not matter: shuffle the deck, rephrase the cue, give a new example, or explain the larger idea without the card.
You are not trying to make the answer mysterious. You are checking whether the knowledge belongs to you, or only to the question that has learned exactly how to ask for it.
References
Al-Najafi, D., Krause, K. D., Wang, Y., et al. (2026). Psychometric performance and student perceptions of AI- versus student-generated multiple-choice questions: A single-center randomized controlled trial. BMC Medical Education.
Barenberg, J., Berse, T., Reimann, L., & Dutke, S. (2021). Testing and transfer: Retrieval practice effects across test formats in English vocabulary learning in school. Applied Cognitive Psychology, 35, 700–710.
Benjamin, A. S., Bjork, R. A., & Schwartz, B. L. (1998). The mismeasure of memory: When retrieval fluency is misleading as a metamnemonic index. Journal of Experimental Psychology: General, 127(1), 55–68.
DiBattista, D., Sinnige-Egger, J.-A., & Fortuna, G. (2014). The "None of the Above" option in multiple-choice testing: An experimental study. The Journal of Experimental Education, 82(2), 168–183.
Eitel, A., Endres, T., & Renkl, A. (2022). Specific questions during retrieval practice are better for texts containing seductive details. Applied Cognitive Psychology, 36(5), 996–1008.
Hinze, S. R., & Wiley, J. (2011). Testing the limits of testing effects using completion tests. Memory, 19(3), 290–304.
Koriat, A., & Bjork, R. A. (2005). Illusions of competence in monitoring one's knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(2), 187–194.
Liu, S., Zheng, Z., Kent, C., & Briscoe, J. (2022). Progressive retrieval practice leads to greater memory for image-word pairs than standard retrieval practice. Memory, 30(7), 796–805.
Lohr, D., Berges, M., Chugh, A., Kohlhase, M., & Müller, D. (2025). Leveraging large language models to generate course-specific semantically annotated learning objects. Journal of Computer Assisted Learning, 41, e13101.
McDaniel, M. A., Thomas, R. C., Agarwal, P. K., McDermott, K. B., & Roediger, H. L. (2013). Quizzing in middle-school science: Successful transfer performance on classroom exams. Applied Cognitive Psychology, 27(3), 360–372.
Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
Smith, M. A., & Karpicke, J. D. (2014). Retrieval practice with short-answer, multiple-choice, and hybrid tests. Memory, 22(7), 784–802.
