The Science Behind Rycal
What the research says about how studying works, how Rycal implements it, and where the evidence stops.
1. The problem: why a tool like Rycal needs to exist
In a landmark review, Dunlosky, Rawson, Marsh, Nathan, and Willingham (2013) evaluated ten common learning techniques. Only two earned high-utility ratings across materials, learners, and contexts. Those two were distributed practice (spacing study over time) and practice testing (retrieval practice). Highlighting and rereading, the two techniques students use most, were rated low utility.
Yet students overwhelmingly use the least effective strategies. Kornell and Bjork (2007) documented that most students study by rereading, and Karpicke, Butler, and Roediger (2009) found that students do not recognize the benefits of testing. When asked, they predict that restudying will produce better long-term retention than retrieval practice, the opposite of what the data show. Koriat and Bjork (2005) described the broader pattern as illusions of competence. The fluency of rereading feels like learning, so students stop studying material they cannot actually recall.
The default student workflow is therefore to reread notes until they feel familiar, then cram the night before the exam. Cramming (massed practice) can work for a test tomorrow (Rohrer, 2006), but the advantage of spacing grows with the retention interval, and for durable learning the gap is large. The distance between what the science recommends and what students actually do is the reason a tool like Rycal exists. It makes the effective strategies the default rather than relying on students to choose them against their own metacognitive instincts.
2. How study tools fail: a taxonomy
Not all study apps implement the science they advertise. The common failure modes, described without naming products:
Passive rereading dressed as flashcards
A card that shows the term and definition together, or that a student flips without attempting recall first, is rereading with extra steps. The testing effect requires a genuine retrieval attempt (Roediger & Karpicke, 2006).
Recognition-only practice
Matching games and pick-the-answer drills test recognition, not recall. Kang, McDermott, and Roediger (2007) found that with feedback provided, short-answer (recall) tests produced better delayed retention than multiple-choice (recognition) tests, consistent with the broader finding that more demanding retrieval produces greater benefit.
Unspaced drilling
Practice without a scheduling mechanism is massed practice by default. Cepeda et al. (2006) synthesized 317 experiments and found spacing beats massing, with the optimal gap scaling to the retention interval (Cepeda et al., 2008).
Testing without feedback
Rowland's (2014) meta-analysis of 61 studies found the testing effect roughly doubles with feedback (g = 0.73) compared to without (g = 0.39). Butler and Roediger (2008) showed that multiple-choice lures can plant false memories that persist to a delayed test; feedback, especially on the distractors, is what prevents testing from backfiring.
Retrieval made too easy
Bjork's framework of desirable difficulties (Bjork, 1994; Bjork & Bjork, 2011; Soderstrom & Bjork, 2015) holds that conditions impairing short-term performance, including effortful retrieval, enhance long-term retention. Karpicke and Roediger (2007) found that equally spaced retrieval, which is harder because more forgetting has occurred, beats expanding schedules for long-term retention.
Gamification as the motivational strategy
Gamification refers to the use of game design elements, such as points, coins, streaks, and leaderboards, in non-game contexts. These are engagement mechanics, not learning techniques. Deci, Koestner, and Ryan (1999) meta-analyzed 128 experiments and found tangible expected rewards undermine intrinsic motivation (d = -0.28 to -0.40), while informational feedback enhances it (d = 0.33). Sailer and Homner (2020) found small cognitive effects of gamification (g = 0.49) but non-significant motivational effects in the most rigorous studies. The benefits are front-loaded and fade: Bai et al. (2020) found almost negligible and negative effects for interventions longer than a semester, and Mazeas et al. (2022) found g = 0.42 during the intervention but g = 0.09, non-significant, at follow-up across 16 randomized controlled trials. Silverman and Barasch (2023) found that people adopt streak maintenance as the goal itself, displacing the original goal, and that broken streaks are especially demotivating; Lally et al. (2010) showed that missing a single day does not affect habit formation, which means streaks punish precisely what does not matter. Leaderboards concentrate the damage on the students who need the most help (Chen et al., 2024; Philpott and Son, 2022). Hanus and Fox (2015) ran the cleanest classroom test and found the gamified course produced declining motivation and lower exam scores. Mekler et al. (2017) found that points and leaderboards drive performance quantity, not quality.
The meta-analyses do show small positive cognitive effects of gamification, and game elements can support onboarding and short-term engagement. What the evidence does not support is treating points, streaks, or leaderboards as a learning strategy. They do not improve retention, their effects expire before a distant exam, and controlling reward structures can undermine the intrinsic motivation that sustains months of study. This distinction is developed further in section 5.
3. The science Rycal is built on
Two strategies earned high-utility ratings from Dunlosky et al. (2013), generalizing across ages, abilities, materials, and criterion tasks. Everything in Rycal is organized around them.
Distributed practice
Cepeda et al. (2006) synthesized 839 assessments from 317 experiments. Spacing study over time beats massing it. Cepeda et al. (2008), with over 1,350 participants, mapped the temporal ridgeline. The optimal gap between study sessions grows with the retention interval, roughly 10 to 20 percent of the delay before the test. Donoghue and Hattie (2021) re-confirmed the finding across 242 studies and 169,179 participants. One boundary condition matters: Rohrer and Taylor (2006) found no spacing benefit at a one-week test but an extremely large benefit at four weeks.
Practice testing
Roediger and Karpicke (2006) had students either restudy a passage four times or read once and recall it three times. At five minutes the restudiers led; at one week the testers retained 61 percent versus 40 percent. Rowland (2014) meta-analyzed 61 studies (g = 0.50 overall), and Adesope, Trevisan, and Sundararajan (2017) confirmed the effect in a large review. Karpicke and Roediger (2008) published the mechanism argument in Science. It is the retrieval itself, not extra exposure, that drives the benefit.
Three elaborations of the testing effect matter for Rycal's design:
Retrieval beats elaboration
Karpicke and Blunt (2011), also in Science, had students either study with concept mapping or practice free recall. Free recall won by about 50 percent at one week, including on inference questions, and even when the final test was itself concept mapping. Blunt and Karpicke (2014) later showed concept mapping only helps when done without materials present, that is, when it becomes retrieval practice.
Feedback is load-bearing
Rowland's (2014) moderator analysis (g = 0.73 with feedback versus 0.39 without) and Butler and Roediger (2008), who showed feedback both boosts correct retention and suppresses lure intrusions, make feedback a requirement rather than a nicety. Butler, Karpicke, and Roediger (2007) found delayed feedback beat immediate feedback on a delayed test, though Hays, Kornell, and Bjork (2012) showed immediate feedback matters most right after a failed retrieval attempt.
Successive relearning
Rawson and Dunlosky (2011; Rawson, Dunlosky, & Sciartelli, 2013) showed that the efficient schedule is to learn material to a criterion of about three correct retrievals, then run relearning sessions (each to one correct retrieval) spaced over time. Benefits asymptote after roughly three retrieval events per session. This is the closest published paradigm to a flashcard loop.
Transfer
Butler (2010) found repeated testing beat restudy on new inferential questions, including far transfer to a new domain (d approx 0.99 at one week). Butler, Black-Maier, Raley, and Marsh (2017) showed retrieval with varied examples transfers better than retrieval with repeated examples. But the boundary is real. Van Gog and Sweller (2015) argued the testing effect shrinks as material complexity rises, and Tran, Rohrer, and Pashler (2015) found no benefit for true deductive inference across four experiments. Karpicke and Aue (2015) pushed back, but the honest summary is that retrieval supports inference-question performance under good conditions and the literature is genuinely split at the boundary.
4. How Rycal implements it, feature by feature
Spaced-repetition scheduler
Every flashcard carries its own FSRS state, tracking difficulty, stability, and retrievability. Rating a card Know or Still Learning updates that state and sets when the card returns. Cards you keep missing come back sooner. Cards you know stay away longer.
FSRS outperformed older algorithms on large benchmarks of recall prediction (Ye et al., 2022), which is why Rycal uses it. That does not make it scientifically optimal. No scheduler has that status, and Rycal has not run its own efficacy trial comparing scheduling methods. The schedule-shape caveats still apply: Karpicke and Roediger (2007) found equally spaced retrieval superior for long-term retention, and Latimier, Peyre, and Ramus (2021) concluded schedule shape matters less than practitioner guidance suggests.
The scheduler is deadline-aware. Intervals are capped so nothing is scheduled past the test date, consistent with Cepeda's finding that optimal spacing scales with the retention interval. Within two days of a test it switches to a triage mode that prioritizes the weakest material. That triage is pragmatic test preparation, not a studied technique, and it is described that way in the product.
Binary grading (Know versus Still Learning) is a deliberate simplification. The FSRS team reports slightly better scheduling accuracy for users who grade with two options than for users who use four, so the simplification is defensible, though this comes from industry research rather than peer-reviewed trials. "Last grade wins" has no direct literature behind it and is not presented as a finding.
Flashcards as retrieval events
The core loop, attempt recall before flipping, then rate honestly, is the testing effect in its canonical form. Covert retrieval works as well as overt retrieval (Smith, Roediger, & Karpicke, 2013), so the memory benefit does not depend on any particular card interface; what matters is the genuine attempt.
Learn: recognition to recall
Learn runs three stages. First pick the definition (recognition), then type the term (cued recall), then write what you remember and self-rate (free recall). Each stage is individually supported, and Kang et al. (2007) directly support the ordering: with feedback, demanding recall tests beat recognition tests on delayed retention. Successive relearning (Rawson & Dunlosky) supports the retry-to-criterion logic for missed items. But no study validates this exact three-stage sequence as a package. It is a well-reasoned scaffold, described as such.
Brain Dump
Free recall is the testing effect in its original form (Roediger & Karpicke, 2006, used free recall as the test condition), and Karpicke and Blunt (2011) is the best single citation for the mode: writing what you remember beats elaborate study activity, including on inference questions.
"Teach it" asks students to explain a concept in writing. Two cautions apply. First, the frequently cited study (Nestojko, Bui, Kornell, & Bjork, 2014) tested expecting to teach, not writing explanations; it supports the framing, not the feature. Second, the explanation must be produced from memory: Koh et al. (2018) found the benefit depends on closed-book generation. Rycal's mode is closed-book, which is the right implementation.
"Cause & effect" asks why and how questions, which is elaborative interrogation. Dunlosky et al. (2013) rated it moderate utility while noting the evidence base is thin and mostly limited to isolated facts rather than real educational contexts.
Multiple-choice questions with explanations
The MCQ banks are retrieval practice with feedback, and the feedback is doing heavy lifting (Rowland, 2014; Butler and Roediger, 2008). Butler, Godbole, and Marsh (2013) found explanation feedback beats answer-only feedback on transfer, which is why Rycal's explanations address the misconception behind each distractor rather than just naming the answer. One safety note from the literature: plausible distractors can plant false memories (Roediger & Marsh, 2005), which makes feedback safety-critical rather than optional. Rycal gives immediate feedback; the literature is genuinely mixed on timing (Butler et al., 2007, favor delayed; Hays et al., 2012, favor immediate after failed retrieval), so immediacy is presented as a design choice, not as the scientifically optimal timing.
FRQ drills
Scaffolded free-response writing, analyze the stimulus, plan, write, self-check against a rubric, sits at the intersection of retrieval practice (strong) and writing-to-learn (moderate). The relevant principle is transfer-appropriate processing (Morris, Bransford, & Franks, 1977): performance is best when the cognitive operations during study match those required at test. An exam that requires writing under time is best prepared for by writing under time. The scaffolding sequence itself is good pedagogy, not a directly validated protocol.
Confidence calibration
After a Brain Dump, Rycal asks students to predict their performance and then shows the measured result. Retrospective confidence after a retrieval attempt is more accurate than prospective judgments (Dougherty et al., 2005, 2018), so this is the better of the two designs, and immediate judgments are the low-accuracy regime (Nelson & Dunlosky, 1991; Dunlosky & Nelson, 1992). But nothing in the literature shows that the act of predicting boosts learning itself. Calibration is metacognitive scaffolding with modest support. It can improve study decisions (Thiede et al., 2003). It is not a proven learning technique.
Deep Work, Test Planner, mastery heatmap
These are productivity and self-regulation tools, not learning-science techniques in the Dunlosky sense. Deep Work deserves a more precise account than "productivity." Its purpose is not the timer. It is the enforced absence of multitasking and task-switching during study. The attention literature supports that specific thing: divided attention during encoding impairs learning (Ophir, Nass, & Wagner, 2009), and switching tasks leaves attention residue that degrades focus on the next task (Leroy, 2009). Planning a specific when and where for study also reliably increases follow-through (Gollwitzer & Sheeran, 2006). Deep Work is scaffolding for the science-backed behaviors: it does not make each retrieval event more effective, it makes sustained, undistracted retrieval practice happen. The Test Planner turns a test date into a scheduling horizon; the heatmap visualizes card states so students can direct effort. The heatmap's fragile/building/stable buckets are product categories, not validated constructs, and the ambient sound in Deep Work is a preference feature, not an evidence-based intervention.
Cram Sheet
The Cram Sheet is a condensed unit review built for the last days before a test. It has three parts. Hidden key terms present sentences with critical terms blanked out. Confusing pairs show commonly mixed-up concepts in comparison tables. Quick definitions show a term with its definition hidden. Tapping reveals the hidden content. The design question was how to make a tap-to-reveal sheet produce retrieval rather than recognition, since the testing effect depends on the attempt, not the exposure.
The sheet requires an attempt before it counts. Students are instructed to say the answer in their head before tapping, and after revealing they must rate whether they knew it. Covert retrieval produces the same retention benefit as overt retrieval (Smith, Roediger, & Karpicke, 2013), so the silent attempt is sufficient. The knew-or-didn't rating serves two purposes. It forces the metacognitive judgment that makes the attempt honest, and it drives the miss queue: items rated as unknown return for a second pass, up to three rounds, following the successive-relearning logic (Rawson, Dunlosky, & Sciartelli, 2013). The session ends with a recap listing what is still shaky, with correct answers visible, so the last encoding is accurate.
The confusing-pairs tables address a specific failure mode of cramming: students can recall each concept in isolation but confuse them under test conditions. Interleaving confusable categories during practice improves discrimination substantially. Rohrer, Dedrick, and Stershic (2015) found interleaved practice outperformed blocked practice with an effect size of d = 0.79 at a 30-day delay, and Taylor and Rohrer (2010) found the same pattern with spacing controlled. The mechanism is discriminative contrast: seeing similar concepts side by side forces attention to the features that distinguish them. The tables put concepts in columns and attributes in rows, with the key-difference row highlighted, which is the contrast made explicit. Pairs are shuffled rather than blocked, consistent with the interleaving finding.
Quick Cram mode limits the sheet to about fifteen key terms in roughly ten minutes. The selection prioritizes confusing pairs first, then cards tagged as required by the course framework, then the student's weakest cards by scheduler difficulty. This is triage logic, not a studied technique. The honest account is that a thirty-minute review sheet is a contradiction: cramming works when it is focused retrieval on high-value material, and the mode exists to enforce that focus. No study validates this exact selection rule.
Sequenced study plan
The sequenced study plan connects the Test Planner, the scheduler, Brain Dump, and practice questions into a single guided session. A student enters a test date, the material covered, and time available. Rycal returns an ordered sequence, typically Brain Dump first, then due-card review, then weakest concepts, then AP-style questions, with time estimates for each step. The next day's plan adjusts based on the previous session's performance, shifting time toward modes and material where the student struggled.
Each component is individually supported. The sequencing itself is product design, not a validated protocol. No study has tested this exact combination of modes in this order, and the paper does not claim otherwise. The rationale is practical rather than empirical: students preparing for a test face a real decision problem, what to study, in what order, for how long, and when to stop. Existing tools answer the components separately. The plan answers the decision. Whether that integration produces better outcomes than students assembling their own sequences is an open empirical question, and Rycal has not run the trial.
5. Motivation without gamification
Gamification, the application of game design elements such as points, coins, streaks, and leaderboards to non-game contexts, is the default motivational strategy of most study apps. These are engagement mechanics, not learning techniques. Rycal uses none of them. The reasoning is motivational, and it cuts against the default design of study apps, so it deserves its own section.
The theoretical frame is self-determination theory (Deci & Ryan): motivation is most durable when it supports autonomy, competence, and relatedness. Deci, Koestner, and Ryan (1999) meta-analyzed 128 experiments and found tangible expected rewards undermine intrinsic motivation (d = -0.28 to -0.40), while informational feedback, feedback that tells you how you are doing without controlling you, enhances it (d = 0.33). The line Rycal draws is exactly that one: informational game elements (progress feedback, mastery visualization, which is what the heatmap is) are compatible with the evidence; controlling ones (coins, loss-aversion streaks, public rank) are not.
The applied evidence is consistent. Sailer and Homner (2020) found small cognitive effects of gamification (g = 0.49) but non-significant motivational effects in the most rigorous studies. The benefits expire: Bai et al. (2020) found almost negligible and negative effects for interventions longer than a semester, and Mazeas et al. (2022) found g = 0.42 during the intervention but g = 0.09, non-significant, at follow-up across 16 randomized controlled trials. For an exam months away, a motivational effect that expires in weeks is the wrong tool.
Streaks deserve specific attention because they are the most common mechanic. Silverman and Barasch (2023) found that people adopt streak maintenance as the goal itself, displacing the original goal, and that broken streaks are especially demotivating. Lally et al. (2010) showed that missing a single day does not affect habit formation, which means streaks punish precisely what does not matter. Leaderboards concentrate the damage on the students who need the most help: Chen et al. (2024) studied the demoralization effect of rank feedback on low performers in a randomized trial, and Philpott and Son (2022) found effort stopping at reward thresholds. Hanus and Fox (2015) ran the cleanest classroom test and found the gamified course produced declining motivation and lower exam scores.
What the literature supports instead is mastery goals over performance goals (Dweck), autonomy, and informational competence feedback (Deci & Ryan). Rycal's motivational design is therefore subtractive: remove the mechanics that redirect goals toward the app, and let the feedback, the heatmap, the review queue, tell the student how the learning itself is going.
This paper does not claim gamification harms learning. The meta-analyses show small positive cognitive effects, and the claim is narrower: gamification does not improve retention, its effects fade, and controlling reward structures can undermine intrinsic motivation. It does not claim streaks never motivate anyone; the claim is about what they optimize for and what breaks. No randomized trial has tested Rycal head-to-head against a gamified competitor, and none is claimed.
6. What Rycal does not claim
A methodology paper should say what is out of bounds, so this section is explicit:
Rycal does not claim its scheduler is scientifically optimal. FSRS is a well-validated implementation of the spacing principle, not a proven optimum, and Rycal has run no trial of its own comparing scheduling methods.
Rycal does not claim Learn's three-stage sequence is an empirically validated protocol. The stages are supported; the sequence is design.
Rycal does not claim the Cram Sheet's exact combination of tap-to-reveal, knew-or-didn't ratings, miss re-queueing, and comparison tables has been validated as a package. The components are supported. The Quick Cram selection rule is triage logic, not a studied technique.
Rycal does not claim the sequenced study plan produces better outcomes than students assembling their own study sequences. No such trial exists.
Rycal does not claim its methods guarantee transfer to AP-style reasoning. The transfer literature is genuinely split, and this paper cites both sides.
Rycal does not claim the mastery heatmap's buckets, the Deep Work timer, or ambient sound are evidence-based interventions. They are product design.
Rycal does not claim gamification harms learning, that streaks never motivate, or that it has been proven superior to any specific product. No such trial exists.
Rycal does not claim specific effect sizes for its own product ("3x better retention"). Effect sizes in the literature depend on retention interval, materials, and comparison conditions, and lab effect sizes do not transfer automatically to a product.
7. References
Adesope, O. O., Trevisan, D. A., & Sundararajan, N. (2017). Rethinking the use of tests: A meta-analysis of practice testing. Review of Educational Research, 87(3), 659-701. https://doi.org/10.3102/0034654316689306
Bai, S., Hew, K. F., & Huang, B. (2020). Does gamification improve student learning outcome? Evidence from a meta-analysis and synthesis of qualitative data in educational contexts. Educational Research Review, 30, 100322. https://doi.org/10.1016/j.edurev.2020.100322
Bertsch, S., Pesta, B. J., Wiscott, R., & McDaniel, M. A. (2007). The generation effect: A meta-analytic review. Memory & Cognition, 35(2), 201-210. https://doi.org/10.3758/BF03193441
Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. Shimamura (Eds.), Metacognition: Knowing about knowing (pp. 185-205). MIT Press.
Bjork, R. A., & Bjork, E. L. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M. A. Gernsbacher et al. (Eds.), Psychology and the real world (pp. 56-64). Worth.
Blunt, J. R., & Karpicke, J. D. (2014). Learning with retrieval-based concept mapping. Journal of Educational Psychology, 106(3), 849-858. https://doi.org/10.1037/a0035934
Butler, A. C. (2010). Repeated testing produces superior transfer of learning relative to repeated studying. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(5), 1118-1133. https://doi.org/10.1037/a0019902
Butler, A. C., Black-Maier, A. C., Raley, N. D., & Marsh, E. J. (2017). Retrieving and applying knowledge to different examples promotes transfer of learning. Journal of Experimental Psychology: Applied, 23(4), 433-446. https://doi.org/10.1037/xap0000142
Butler, A. C., Godbole, N., & Marsh, E. J. (2013). Explanation feedback is better than correct answer feedback for promoting transfer of learning. Journal of Educational Psychology, 105(2), 290-298. https://doi.org/10.1037/a0031026
Butler, A. C., Karpicke, J. D., & Roediger, H. L., III. (2007). The effect of type and timing of feedback on learning from multiple-choice tests. Journal of Experimental Psychology: Applied, 13(4), 273-281. https://doi.org/10.1037/1076-898X.13.4.273
Butler, A. C., & Roediger, H. L., III. (2008). Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing. Memory & Cognition, 36(3), 604-616. https://doi.org/10.3758/MC.36.3.604
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354-380. https://doi.org/10.1037/0033-2909.132.3.354
Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1124-1133. https://doi.org/10.1111/j.1467-9280.2008.02209.x
Chen, J., Dobrescu, L. I., Foster, G., & Motta, A. (2024). Can leagues mitigate the demoralization effect of rank feedback? A randomized controlled trial. Labour Economics, 90, 102602. https://doi.org/10.1016/j.labeco.2024.102602
Chi, M. T. H., Bassok, M., Lewis, M. W., Reimann, P., & Glaser, R. (1989). Self-explanations: How students study and use examples in learning to solve problems. Cognitive Science, 13(2), 145-182. https://doi.org/10.1207/s15516709cog1302_1
Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627-668. https://doi.org/10.1037/0033-2909.125.6.627
Ryan, R. M., & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68-78. https://doi.org/10.1037/0003-066X.55.1.68
Dweck, C. S. (2006). Mindset: The new psychology of success. Random House.
Dougherty, M. R., Scheck, P., Nelson, T. O., & Narens, L. (2005). Using the past to predict the future. Memory & Cognition, 33(6), 1096-1115. https://doi.org/10.3758/BF03193216
Dougherty, M. R., Robey, A. M., & Buttaccio, D. R. (2018). Do metacognitive judgments alter memory performance beyond the benefits of retrieval practice? Memory & Cognition, 46(4), 558-565. https://doi.org/10.3758/s13421-017-0778-7
Dunlosky, J., & Nelson, T. O. (1992). Importance of the kind of cue for judgments of learning (JOL) and the delayed-JOL effect. Memory & Cognition, 20(4), 374-380. https://doi.org/10.3758/BF03210921
Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving Students' Learning With Effective Learning Techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest, 14(1), 4-58. https://doi.org/10.1177/1529100612453266
Ebersbach, M., Feierabend, M., & Barzagar Nazari, K. (2020). Comparing the effects of generating questions, testing, and restudying on students' long-term recall in university learning. Applied Cognitive Psychology, 34(4), 724-740. https://doi.org/10.1002/acp.3639
Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. Advances in Experimental Social Psychology, 38, 69-119. https://doi.org/10.1016/S0065-2601(06)38002-1
Hanus, M. D., & Fox, J. (2015). Assessing the effects of gamification in the classroom: A longitudinal study on intrinsic motivation, social comparison, satisfaction, effort, and academic performance. Computers & Education, 80, 152-161. https://doi.org/10.1016/j.compedu.2014.08.019
Donoghue, G. M., & Hattie, J. A. C. (2021). A meta-analysis of ten learning techniques. Frontiers in Education, 6, 581216. https://doi.org/10.3389/feduc.2021.581216
Hays, M. J., Kornell, N., & Bjork, R. A. (2013). When and why a failed test potentiates the effectiveness of subsequent study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 39(1), 290-296. https://doi.org/10.1037/a0028468
Kang, S. H. K., McDermott, K. B., & Roediger, H. L., III. (2007). Test format and corrective feedback modify the effect of testing on long-term retention. European Journal of Cognitive Psychology, 19(4-5), 528-558. https://doi.org/10.1080/09541440601056620
Karpicke, J. D., & Aue, W. R. (2015). The testing effect is alive and well with complex materials. Educational Psychology Review, 27(2), 317-326. https://doi.org/10.1007/s10648-015-9309-3
Karpicke, J. D., & Blunt, J. R. (2011). Retrieval practice produces more learning than elaborative studying with concept mapping. Science, 331(6018), 772-775. https://doi.org/10.1126/science.1199327
Karpicke, J. D., Butler, A. C., & Roediger, H. L., III. (2009). Metacognitive strategies in student learning: Do students practise retrieval when they study on their own? Memory, 17(4), 471-479. https://doi.org/10.1080/09658210802647009
Karpicke, J. D., & Roediger, H. L., III. (2007). Expanding retrieval practice promotes short-term retention, but equally spaced retrieval enhances long-term retention. Journal of Experimental Psychology: Learning, Memory, and Cognition, 33(4), 704-719. https://doi.org/10.1037/0278-7393.33.4.704
Karpicke, J. D., & Roediger, H. L., III. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966-968. https://doi.org/10.1126/science.1152408
Kimball, D. R., & Metcalfe, J. (2003). Delaying judgments of learning affects memory, not metamemory. Memory & Cognition, 31(6), 918-929. https://doi.org/10.3758/BF03196445
Koh, A. W. L., Lee, S. C., & Lim, S. W. H. (2018). The learning benefits of teaching: A retrieval practice hypothesis. Applied Cognitive Psychology, 32(3), 401-410. https://doi.org/10.1002/acp.3410
Koriat, A., & Bjork, R. A. (2005). Illusions of competence in monitoring one's knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(2), 187-194. https://doi.org/10.1037/0278-7393.31.2.187
Kornell, N., & Bjork, R. A. (2007). The promise and perils of self-regulated study. Psychonomic Bulletin & Review, 14(2), 219-224. https://doi.org/10.3758/BF03194055
Leroy, S. (2009). Why is it so hard to do my work? The challenge of attention residue when switching between work tasks. Organizational Behavior and Human Decision Processes, 109(2), 168-181. https://doi.org/10.1016/j.obhdp.2009.04.002
Kornell, N., Hays, M. J., & Bjork, R. A. (2009). Unsuccessful retrieval attempts enhance subsequent learning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(4), 989-998. https://doi.org/10.1037/a0015729
Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. European Journal of Social Psychology, 40(6), 998-1009. https://doi.org/10.1002/ejsp.674
Landauer, T. K., & Bjork, R. A. (1978). Optimum rehearsal patterns and name learning. In M. M. Gruneberg et al. (Eds.), Practical aspects of memory (pp. 625-632). Academic Press.
Latimier, A., Peyre, H., & Ramus, F. (2021). A meta-analytic review of the benefit of spacing out retrieval practice episodes on retention. Educational Psychology Review, 33, 959-987. https://doi.org/10.1007/s10648-020-09572-8
Mazeas, A., Duclos, M., Pereira, B., & Chalabaev, A. (2022). Evaluating the effectiveness of gamification on physical activity: Systematic review and meta-analysis of randomized controlled trials. Journal of Medical Internet Research, 24(1), e26779. https://doi.org/10.2196/26779
Mekler, E. D., Brühlmann, F., Tuch, A. N., & Opwis, K. (2017). Towards understanding the effects of individual gamification elements on intrinsic motivation and performance. Computers in Human Behavior, 71, 525-534. https://doi.org/10.1016/j.chb.2015.08.048
Morris, C. D., Bransford, J. D., & Franks, J. J. (1977). Levels of processing versus transfer appropriate processing. Journal of Verbal Learning and Verbal Behavior, 16(5), 519-533. https://doi.org/10.1016/S0022-5371(77)80016-9
Nelson, T. O., & Dunlosky, J. (1991). When people's judgments of learning (JOLs) are extremely accurate at predicting subsequent recall: The "delayed-JOL effect." Psychological Science, 2(4), 267-270. https://doi.org/10.1111/j.1467-9280.1991.tb00147.x
Nestojko, J. F., Bui, D. C., Kornell, N., & Bjork, E. L. (2014). Expecting to teach enhances learning and organization of knowledge in free recall of text passages. Memory & Cognition, 42(7), 1038-1048. https://doi.org/10.3758/s13421-014-0416-z
Ophir, E., Nass, C., & Wagner, A. D. (2009). Cognitive control in media multitaskers. Proceedings of the National Academy of Sciences, 106(37), 15583-15587. https://doi.org/10.1073/pnas.0903620106
Philpott, A., & Son, J.-B. (2022). Leaderboards in an EFL course: Student performance and motivation. Computers & Education, 190, 104605. https://doi.org/10.1016/j.compedu.2022.104605
Rawson, K. A., & Dunlosky, J. (2011). Optimizing schedules of retrieval practice for durable and efficient learning: How much is enough? Journal of Experimental Psychology: General, 140(3), 283-302. https://doi.org/10.1037/a0022547
Rawson, K. A., Dunlosky, J., & Sciartelli, S. M. (2013). The power of successive relearning: Improving performance on course exams and long-term retention. Educational Psychology Review, 25(4), 523-548. https://doi.org/10.1007/s10648-013-9240-4
Richland, L. E., Kornell, N., & Kao, L. S. (2009). The pretesting effect: Do unsuccessful retrieval attempts enhance learning? Journal of Experimental Psychology: Applied, 15(3), 243-257. https://doi.org/10.1037/a0016496
Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
Roediger, H. L., III, & Marsh, E. J. (2005). The positive and negative consequences of multiple-choice testing. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(5), 1155-1159. https://doi.org/10.1037/0278-7393.31.5.1155
Rohrer, D., Dedrick, R. F., & Stershic, S. (2015). Interleaved practice improves mathematics learning. Journal of Educational Psychology, 107(3), 900-908. https://doi.org/10.1037/edu0000001
Rohrer, D., & Taylor, K. (2006). The effects of overlearning and distributed practice on the retention of mathematics knowledge. Applied Cognitive Psychology, 20(9), 1209-1224. https://doi.org/10.1002/acp.1266
Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432-1463. https://doi.org/10.1037/a0037559
Taylor, K., & Rohrer, D. (2010). The effects of interleaved practice. Applied Cognitive Psychology, 24(6), 837-848. https://doi.org/10.1002/acp.1598
Sailer, M., & Homner, L. (2020). The gamification of learning: A meta-analysis. Educational Psychology Review, 32(1), 77-112. https://doi.org/10.1007/s10648-019-09498-w
Silverman, J., & Barasch, A. (2023). On or off track: How (broken) streaks affect consumer decisions. Journal of Consumer Research, 49(6), 1095-1117. https://doi.org/10.1093/jcr/ucac029
Slamecka, N. J., & Graf, P. (1978). The generation effect: Delineation of a phenomenon. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592-604. https://doi.org/10.1037/0278-7393.4.6.592
Smith, M. A., Roediger, H. L., III, & Karpicke, J. D. (2013). Covert retrieval practice benefits retention as much as overt retrieval practice. Journal of Experimental Psychology: Learning, Memory, and Cognition, 39(6), 1712-1725. https://doi.org/10.1037/a0033569
Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176-199. https://doi.org/10.1177/1745691615569000
Thiede, K. W., Anderson, M. C. M., & Therriault, D. J. (2003). Accuracy of metacognitive monitoring affects learning of texts. Journal of Educational Psychology, 95(1), 66-73. https://doi.org/10.1037/0022-0663.95.1.66
Tran, R., Rohrer, D., & Pashler, H. (2015). Retrieval practice: The lack of transfer to deductive inferences. Psychonomic Bulletin & Review, 22(1), 135-140. https://doi.org/10.3758/s13423-014-0646-x
van Gog, T., & Sweller, J. (2015). Not new, but nearly forgotten: The testing effect decreases or even disappears as the complexity of learning materials increases. Educational Psychology Review, 27(2), 247-264. https://doi.org/10.1007/s10648-015-9310-x
Ye, J., Su, J., & Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4381-4390. https://doi.org/10.1145/3534678.3539081