Rycal Open the app
Methodology white paper

The Science Behind Rycal

What the research says about how studying works, how Rycal implements it, and where the evidence stops.

Rycal · October 2026 · Full references at the end

1. The problem: why a tool like Rycal needs to exist

In a landmark review, Dunlosky, Rawson, Marsh, Nathan, and Willingham (2013) evaluated ten common learning techniques. Only two earned high-utility ratings across materials, learners, and contexts. Those two were distributed practice (spacing study over time) and practice testing (retrieval practice). Highlighting and rereading, the two techniques students use most, were rated low utility.

Yet students overwhelmingly use the least effective strategies. Kornell and Bjork (2007) documented that most students study by rereading, and Karpicke, Butler, and Roediger (2009) found that students do not recognize the benefits of testing. When asked, they predict that restudying will produce better long-term retention than retrieval practice, the opposite of what the data show. Koriat and Bjork (2005) described the broader pattern as illusions of competence. The fluency of rereading feels like learning, so students stop studying material they cannot actually recall.

The default student workflow is therefore to reread notes until they feel familiar, then cram the night before the exam. Cramming (massed practice) can work for a test tomorrow (Rohrer, 2006), but the advantage of spacing grows with the retention interval, and for durable learning the gap is large. The distance between what the science recommends and what students actually do is the reason a tool like Rycal exists. It makes the effective strategies the default rather than relying on students to choose them against their own metacognitive instincts.

2. How study tools fail: a taxonomy

Not all study apps implement the science they advertise. The common failure modes, described without naming products:

Passive rereading dressed as flashcards

A card that shows the term and definition together, or that a student flips without attempting recall first, is rereading with extra steps. The testing effect requires a genuine retrieval attempt (Roediger & Karpicke, 2006).

Recognition-only practice

Matching games and pick-the-answer drills test recognition, not recall. Kang, McDermott, and Roediger (2007) found that with feedback provided, short-answer (recall) tests produced better delayed retention than multiple-choice (recognition) tests, consistent with the broader finding that more demanding retrieval produces greater benefit.

Unspaced drilling

Practice without a scheduling mechanism is massed practice by default. Cepeda et al. (2006) synthesized 317 experiments and found spacing beats massing, with the optimal gap scaling to the retention interval (Cepeda et al., 2008).

Testing without feedback

Rowland's (2014) meta-analysis of 61 studies found the testing effect roughly doubles with feedback (g = 0.73) compared to without (g = 0.39). Butler and Roediger (2008) showed that multiple-choice lures can plant false memories that persist to a delayed test; feedback, especially on the distractors, is what prevents testing from backfiring.

Retrieval made too easy

Bjork's framework of desirable difficulties (Bjork, 1994; Bjork & Bjork, 2011; Soderstrom & Bjork, 2015) holds that conditions impairing short-term performance, including effortful retrieval, enhance long-term retention. Karpicke and Roediger (2007) found that equally spaced retrieval, which is harder because more forgetting has occurred, beats expanding schedules for long-term retention.

Gamification as the motivational strategy

Gamification refers to the use of game design elements, such as points, coins, streaks, and leaderboards, in non-game contexts. These are engagement mechanics, not learning techniques. Deci, Koestner, and Ryan (1999) meta-analyzed 128 experiments and found tangible expected rewards undermine intrinsic motivation (d = -0.28 to -0.40), while informational feedback enhances it (d = 0.33). Sailer and Homner (2020) found small cognitive effects of gamification (g = 0.49) but non-significant motivational effects in the most rigorous studies. The benefits are front-loaded and fade: Bai et al. (2020) found almost negligible and negative effects for interventions longer than a semester, and Mazeas et al. (2022) found g = 0.42 during the intervention but g = 0.09, non-significant, at follow-up across 16 randomized controlled trials. Silverman and Barasch (2023) found that people adopt streak maintenance as the goal itself, displacing the original goal, and that broken streaks are especially demotivating; Lally et al. (2010) showed that missing a single day does not affect habit formation, which means streaks punish precisely what does not matter. Leaderboards concentrate the damage on the students who need the most help (Chen et al., 2024; Philpott and Son, 2022). Hanus and Fox (2015) ran the cleanest classroom test and found the gamified course produced declining motivation and lower exam scores. Mekler et al. (2017) found that points and leaderboards drive performance quantity, not quality.

The meta-analyses do show small positive cognitive effects of gamification, and game elements can support onboarding and short-term engagement. What the evidence does not support is treating points, streaks, or leaderboards as a learning strategy. They do not improve retention, their effects expire before a distant exam, and controlling reward structures can undermine the intrinsic motivation that sustains months of study. This distinction is developed further in section 5.

3. The science Rycal is built on

Two strategies earned high-utility ratings from Dunlosky et al. (2013), generalizing across ages, abilities, materials, and criterion tasks. Everything in Rycal is organized around them.

Distributed practice

Cepeda et al. (2006) synthesized 839 assessments from 317 experiments. Spacing study over time beats massing it. Cepeda et al. (2008), with over 1,350 participants, mapped the temporal ridgeline. The optimal gap between study sessions grows with the retention interval, roughly 10 to 20 percent of the delay before the test. Donoghue and Hattie (2021) re-confirmed the finding across 242 studies and 169,179 participants. One boundary condition matters: Rohrer and Taylor (2006) found no spacing benefit at a one-week test but an extremely large benefit at four weeks.

Practice testing

Roediger and Karpicke (2006) had students either restudy a passage four times or read once and recall it three times. At five minutes the restudiers led; at one week the testers retained 61 percent versus 40 percent. Rowland (2014) meta-analyzed 61 studies (g = 0.50 overall), and Adesope, Trevisan, and Sundararajan (2017) confirmed the effect in a large review. Karpicke and Roediger (2008) published the mechanism argument in Science. It is the retrieval itself, not extra exposure, that drives the benefit.

Three elaborations of the testing effect matter for Rycal's design:

Retrieval beats elaboration

Karpicke and Blunt (2011), also in Science, had students either study with concept mapping or practice free recall. Free recall won by about 50 percent at one week, including on inference questions, and even when the final test was itself concept mapping. Blunt and Karpicke (2014) later showed concept mapping only helps when done without materials present, that is, when it becomes retrieval practice.

Feedback is load-bearing

Rowland's (2014) moderator analysis (g = 0.73 with feedback versus 0.39 without) and Butler and Roediger (2008), who showed feedback both boosts correct retention and suppresses lure intrusions, make feedback a requirement rather than a nicety. Butler, Karpicke, and Roediger (2007) found delayed feedback beat immediate feedback on a delayed test, though Hays, Kornell, and Bjork (2012) showed immediate feedback matters most right after a failed retrieval attempt.

Successive relearning

Rawson and Dunlosky (2011; Rawson, Dunlosky, & Sciartelli, 2013) showed that the efficient schedule is to learn material to a criterion of about three correct retrievals, then run relearning sessions (each to one correct retrieval) spaced over time. Benefits asymptote after roughly three retrieval events per session. This is the closest published paradigm to a flashcard loop.

Transfer

Butler (2010) found repeated testing beat restudy on new inferential questions, including far transfer to a new domain (d approx 0.99 at one week). Butler, Black-Maier, Raley, and Marsh (2017) showed retrieval with varied examples transfers better than retrieval with repeated examples. But the boundary is real. Van Gog and Sweller (2015) argued the testing effect shrinks as material complexity rises, and Tran, Rohrer, and Pashler (2015) found no benefit for true deductive inference across four experiments. Karpicke and Aue (2015) pushed back, but the honest summary is that retrieval supports inference-question performance under good conditions and the literature is genuinely split at the boundary.

4. How Rycal implements it, feature by feature

Spaced-repetition scheduler

Every flashcard carries its own FSRS state, tracking difficulty, stability, and retrievability. Rating a card Know or Still Learning updates that state and sets when the card returns. Cards you keep missing come back sooner. Cards you know stay away longer.

FSRS outperformed older algorithms on large benchmarks of recall prediction (Ye et al., 2022), which is why Rycal uses it. That does not make it scientifically optimal. No scheduler has that status, and Rycal has not run its own efficacy trial comparing scheduling methods. The schedule-shape caveats still apply: Karpicke and Roediger (2007) found equally spaced retrieval superior for long-term retention, and Latimier, Peyre, and Ramus (2021) concluded schedule shape matters less than practitioner guidance suggests.

The scheduler is deadline-aware. Intervals are capped so nothing is scheduled past the test date, consistent with Cepeda's finding that optimal spacing scales with the retention interval. Within two days of a test it switches to a triage mode that prioritizes the weakest material. That triage is pragmatic test preparation, not a studied technique, and it is described that way in the product.

Binary grading (Know versus Still Learning) is a deliberate simplification. The FSRS team reports slightly better scheduling accuracy for users who grade with two options than for users who use four, so the simplification is defensible, though this comes from industry research rather than peer-reviewed trials. "Last grade wins" has no direct literature behind it and is not presented as a finding.

Flashcards as retrieval events

The core loop, attempt recall before flipping, then rate honestly, is the testing effect in its canonical form. Covert retrieval works as well as overt retrieval (Smith, Roediger, & Karpicke, 2013), so the memory benefit does not depend on any particular card interface; what matters is the genuine attempt.

Learn: recognition to recall

Learn runs three stages. First pick the definition (recognition), then type the term (cued recall), then write what you remember and self-rate (free recall). Each stage is individually supported, and Kang et al. (2007) directly support the ordering: with feedback, demanding recall tests beat recognition tests on delayed retention. Successive relearning (Rawson & Dunlosky) supports the retry-to-criterion logic for missed items. But no study validates this exact three-stage sequence as a package. It is a well-reasoned scaffold, described as such.

Brain Dump

Free recall is the testing effect in its original form (Roediger & Karpicke, 2006, used free recall as the test condition), and Karpicke and Blunt (2011) is the best single citation for the mode: writing what you remember beats elaborate study activity, including on inference questions.

"Teach it" asks students to explain a concept in writing. Two cautions apply. First, the frequently cited study (Nestojko, Bui, Kornell, & Bjork, 2014) tested expecting to teach, not writing explanations; it supports the framing, not the feature. Second, the explanation must be produced from memory: Koh et al. (2018) found the benefit depends on closed-book generation. Rycal's mode is closed-book, which is the right implementation.

"Cause & effect" asks why and how questions, which is elaborative interrogation. Dunlosky et al. (2013) rated it moderate utility while noting the evidence base is thin and mostly limited to isolated facts rather than real educational contexts.

Multiple-choice questions with explanations

The MCQ banks are retrieval practice with feedback, and the feedback is doing heavy lifting (Rowland, 2014; Butler and Roediger, 2008). Butler, Godbole, and Marsh (2013) found explanation feedback beats answer-only feedback on transfer, which is why Rycal's explanations address the misconception behind each distractor rather than just naming the answer. One safety note from the literature: plausible distractors can plant false memories (Roediger & Marsh, 2005), which makes feedback safety-critical rather than optional. Rycal gives immediate feedback; the literature is genuinely mixed on timing (Butler et al., 2007, favor delayed; Hays et al., 2012, favor immediate after failed retrieval), so immediacy is presented as a design choice, not as the scientifically optimal timing.

FRQ drills

Scaffolded free-response writing, analyze the stimulus, plan, write, self-check against a rubric, sits at the intersection of retrieval practice (strong) and writing-to-learn (moderate). The relevant principle is transfer-appropriate processing (Morris, Bransford, & Franks, 1977): performance is best when the cognitive operations during study match those required at test. An exam that requires writing under time is best prepared for by writing under time. The scaffolding sequence itself is good pedagogy, not a directly validated protocol.

Confidence calibration

After a Brain Dump, Rycal asks students to predict their performance and then shows the measured result. Retrospective confidence after a retrieval attempt is more accurate than prospective judgments (Dougherty et al., 2005, 2018), so this is the better of the two designs, and immediate judgments are the low-accuracy regime (Nelson & Dunlosky, 1991; Dunlosky & Nelson, 1992). But nothing in the literature shows that the act of predicting boosts learning itself. Calibration is metacognitive scaffolding with modest support. It can improve study decisions (Thiede et al., 2003). It is not a proven learning technique.

Deep Work, Test Planner, mastery heatmap

These are productivity and self-regulation tools, not learning-science techniques in the Dunlosky sense. Deep Work deserves a more precise account than "productivity." Its purpose is not the timer. It is the enforced absence of multitasking and task-switching during study. The attention literature supports that specific thing: divided attention during encoding impairs learning (Ophir, Nass, & Wagner, 2009), and switching tasks leaves attention residue that degrades focus on the next task (Leroy, 2009). Planning a specific when and where for study also reliably increases follow-through (Gollwitzer & Sheeran, 2006). Deep Work is scaffolding for the science-backed behaviors: it does not make each retrieval event more effective, it makes sustained, undistracted retrieval practice happen. The Test Planner turns a test date into a scheduling horizon; the heatmap visualizes card states so students can direct effort. The heatmap's fragile/building/stable buckets are product categories, not validated constructs, and the ambient sound in Deep Work is a preference feature, not an evidence-based intervention.

Cram Sheet

The Cram Sheet is a condensed unit review built for the last days before a test. It has three parts. Hidden key terms present sentences with critical terms blanked out. Confusing pairs show commonly mixed-up concepts in comparison tables. Quick definitions show a term with its definition hidden. Tapping reveals the hidden content. The design question was how to make a tap-to-reveal sheet produce retrieval rather than recognition, since the testing effect depends on the attempt, not the exposure.

The sheet requires an attempt before it counts. Students are instructed to say the answer in their head before tapping, and after revealing they must rate whether they knew it. Covert retrieval produces the same retention benefit as overt retrieval (Smith, Roediger, & Karpicke, 2013), so the silent attempt is sufficient. The knew-or-didn't rating serves two purposes. It forces the metacognitive judgment that makes the attempt honest, and it drives the miss queue: items rated as unknown return for a second pass, up to three rounds, following the successive-relearning logic (Rawson, Dunlosky, & Sciartelli, 2013). The session ends with a recap listing what is still shaky, with correct answers visible, so the last encoding is accurate.

The confusing-pairs tables address a specific failure mode of cramming: students can recall each concept in isolation but confuse them under test conditions. Interleaving confusable categories during practice improves discrimination substantially. Rohrer, Dedrick, and Stershic (2015) found interleaved practice outperformed blocked practice with an effect size of d = 0.79 at a 30-day delay, and Taylor and Rohrer (2010) found the same pattern with spacing controlled. The mechanism is discriminative contrast: seeing similar concepts side by side forces attention to the features that distinguish them. The tables put concepts in columns and attributes in rows, with the key-difference row highlighted, which is the contrast made explicit. Pairs are shuffled rather than blocked, consistent with the interleaving finding.

Quick Cram mode limits the sheet to about fifteen key terms in roughly ten minutes. The selection prioritizes confusing pairs first, then cards tagged as required by the course framework, then the student's weakest cards by scheduler difficulty. This is triage logic, not a studied technique. The honest account is that a thirty-minute review sheet is a contradiction: cramming works when it is focused retrieval on high-value material, and the mode exists to enforce that focus. No study validates this exact selection rule.

Sequenced study plan

The sequenced study plan connects the Test Planner, the scheduler, Brain Dump, and practice questions into a single guided session. A student enters a test date, the material covered, and time available. Rycal returns an ordered sequence, typically Brain Dump first, then due-card review, then weakest concepts, then AP-style questions, with time estimates for each step. The next day's plan adjusts based on the previous session's performance, shifting time toward modes and material where the student struggled.

Each component is individually supported. The sequencing itself is product design, not a validated protocol. No study has tested this exact combination of modes in this order, and the paper does not claim otherwise. The rationale is practical rather than empirical: students preparing for a test face a real decision problem, what to study, in what order, for how long, and when to stop. Existing tools answer the components separately. The plan answers the decision. Whether that integration produces better outcomes than students assembling their own sequences is an open empirical question, and Rycal has not run the trial.

5. Motivation without gamification

Gamification, the application of game design elements such as points, coins, streaks, and leaderboards to non-game contexts, is the default motivational strategy of most study apps. These are engagement mechanics, not learning techniques. Rycal uses none of them. The reasoning is motivational, and it cuts against the default design of study apps, so it deserves its own section.

The theoretical frame is self-determination theory (Deci & Ryan): motivation is most durable when it supports autonomy, competence, and relatedness. Deci, Koestner, and Ryan (1999) meta-analyzed 128 experiments and found tangible expected rewards undermine intrinsic motivation (d = -0.28 to -0.40), while informational feedback, feedback that tells you how you are doing without controlling you, enhances it (d = 0.33). The line Rycal draws is exactly that one: informational game elements (progress feedback, mastery visualization, which is what the heatmap is) are compatible with the evidence; controlling ones (coins, loss-aversion streaks, public rank) are not.

The applied evidence is consistent. Sailer and Homner (2020) found small cognitive effects of gamification (g = 0.49) but non-significant motivational effects in the most rigorous studies. The benefits expire: Bai et al. (2020) found almost negligible and negative effects for interventions longer than a semester, and Mazeas et al. (2022) found g = 0.42 during the intervention but g = 0.09, non-significant, at follow-up across 16 randomized controlled trials. For an exam months away, a motivational effect that expires in weeks is the wrong tool.

Streaks deserve specific attention because they are the most common mechanic. Silverman and Barasch (2023) found that people adopt streak maintenance as the goal itself, displacing the original goal, and that broken streaks are especially demotivating. Lally et al. (2010) showed that missing a single day does not affect habit formation, which means streaks punish precisely what does not matter. Leaderboards concentrate the damage on the students who need the most help: Chen et al. (2024) studied the demoralization effect of rank feedback on low performers in a randomized trial, and Philpott and Son (2022) found effort stopping at reward thresholds. Hanus and Fox (2015) ran the cleanest classroom test and found the gamified course produced declining motivation and lower exam scores.

What the literature supports instead is mastery goals over performance goals (Dweck), autonomy, and informational competence feedback (Deci & Ryan). Rycal's motivational design is therefore subtractive: remove the mechanics that redirect goals toward the app, and let the feedback, the heatmap, the review queue, tell the student how the learning itself is going.

This paper does not claim gamification harms learning. The meta-analyses show small positive cognitive effects, and the claim is narrower: gamification does not improve retention, its effects fade, and controlling reward structures can undermine intrinsic motivation. It does not claim streaks never motivate anyone; the claim is about what they optimize for and what breaks. No randomized trial has tested Rycal head-to-head against a gamified competitor, and none is claimed.

6. What Rycal does not claim

A methodology paper should say what is out of bounds, so this section is explicit:

Rycal does not claim its scheduler is scientifically optimal. FSRS is a well-validated implementation of the spacing principle, not a proven optimum, and Rycal has run no trial of its own comparing scheduling methods.

Rycal does not claim Learn's three-stage sequence is an empirically validated protocol. The stages are supported; the sequence is design.

Rycal does not claim the Cram Sheet's exact combination of tap-to-reveal, knew-or-didn't ratings, miss re-queueing, and comparison tables has been validated as a package. The components are supported. The Quick Cram selection rule is triage logic, not a studied technique.

Rycal does not claim the sequenced study plan produces better outcomes than students assembling their own study sequences. No such trial exists.

Rycal does not claim its methods guarantee transfer to AP-style reasoning. The transfer literature is genuinely split, and this paper cites both sides.

Rycal does not claim the mastery heatmap's buckets, the Deep Work timer, or ambient sound are evidence-based interventions. They are product design.

Rycal does not claim gamification harms learning, that streaks never motivate, or that it has been proven superior to any specific product. No such trial exists.

Rycal does not claim specific effect sizes for its own product ("3x better retention"). Effect sizes in the literature depend on retention interval, materials, and comparison conditions, and lab effect sizes do not transfer automatically to a product.

7. References

Adesope, O. O., Trevisan, D. A., & Sundararajan, N. (2017). Rethinking the use of tests: A meta-analysis of practice testing. Review of Educational Research, 87(3), 659-701. https://doi.org/10.3102/0034654316689306

Bai, S., Hew, K. F., & Huang, B. (2020). Does gamification improve student learning outcome? Evidence from a meta-analysis and synthesis of qualitative data in educational contexts. Educational Research Review, 30, 100322. https://doi.org/10.1016/j.edurev.2020.100322

Bertsch, S., Pesta, B. J., Wiscott, R., & McDaniel, M. A. (2007). The generation effect: A meta-analytic review. Memory & Cognition, 35(2), 201-210. https://doi.org/10.3758/BF03193441

Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. Shimamura (Eds.), Metacognition: Knowing about knowing (pp. 185-205). MIT Press.

Bjork, R. A., & Bjork, E. L. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M. A. Gernsbacher et al. (Eds.), Psychology and the real world (pp. 56-64). Worth.

Blunt, J. R., & Karpicke, J. D. (2014). Learning with retrieval-based concept mapping. Journal of Educational Psychology, 106(3), 849-858. https://doi.org/10.1037/a0035934

Butler, A. C. (2010). Repeated testing produces superior transfer of learning relative to repeated studying. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(5), 1118-1133. https://doi.org/10.1037/a0019902

Butler, A. C., Black-Maier, A. C., Raley, N. D., & Marsh, E. J. (2017). Retrieving and applying knowledge to different examples promotes transfer of learning. Journal of Experimental Psychology: Applied, 23(4), 433-446. https://doi.org/10.1037/xap0000142

Butler, A. C., Godbole, N., & Marsh, E. J. (2013). Explanation feedback is better than correct answer feedback for promoting transfer of learning. Journal of Educational Psychology, 105(2), 290-298. https://doi.org/10.1037/a0031026

Butler, A. C., Karpicke, J. D., & Roediger, H. L., III. (2007). The effect of type and timing of feedback on learning from multiple-choice tests. Journal of Experimental Psychology: Applied, 13(4), 273-281. https://doi.org/10.1037/1076-898X.13.4.273

Butler, A. C., & Roediger, H. L., III. (2008). Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing. Memory & Cognition, 36(3), 604-616. https://doi.org/10.3758/MC.36.3.604

Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354-380. https://doi.org/10.1037/0033-2909.132.3.354

Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1124-1133. https://doi.org/10.1111/j.1467-9280.2008.02209.x

Chen, J., Dobrescu, L. I., Foster, G., & Motta, A. (2024). Can leagues mitigate the demoralization effect of rank feedback? A randomized controlled trial. Labour Economics, 90, 102602. https://doi.org/10.1016/j.labeco.2024.102602

Chi, M. T. H., Bassok, M., Lewis, M. W., Reimann, P., & Glaser, R. (1989). Self-explanations: How students study and use examples in learning to solve problems. Cognitive Science, 13(2), 145-182. https://doi.org/10.1207/s15516709cog1302_1

Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627-668. https://doi.org/10.1037/0033-2909.125.6.627

Ryan, R. M., & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68-78. https://doi.org/10.1037/0003-066X.55.1.68

Dweck, C. S. (2006). Mindset: The new psychology of success. Random House.

Dougherty, M. R., Scheck, P., Nelson, T. O., & Narens, L. (2005). Using the past to predict the future. Memory & Cognition, 33(6), 1096-1115. https://doi.org/10.3758/BF03193216

Dougherty, M. R., Robey, A. M., & Buttaccio, D. R. (2018). Do metacognitive judgments alter memory performance beyond the benefits of retrieval practice? Memory & Cognition, 46(4), 558-565. https://doi.org/10.3758/s13421-017-0778-7

Dunlosky, J., & Nelson, T. O. (1992). Importance of the kind of cue for judgments of learning (JOL) and the delayed-JOL effect. Memory & Cognition, 20(4), 374-380. https://doi.org/10.3758/BF03210921

Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving Students' Learning With Effective Learning Techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest, 14(1), 4-58. https://doi.org/10.1177/1529100612453266

Ebersbach, M., Feierabend, M., & Barzagar Nazari, K. (2020). Comparing the effects of generating questions, testing, and restudying on students' long-term recall in university learning. Applied Cognitive Psychology, 34(4), 724-740. https://doi.org/10.1002/acp.3639

Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. Advances in Experimental Social Psychology, 38, 69-119. https://doi.org/10.1016/S0065-2601(06)38002-1

Hanus, M. D., & Fox, J. (2015). Assessing the effects of gamification in the classroom: A longitudinal study on intrinsic motivation, social comparison, satisfaction, effort, and academic performance. Computers & Education, 80, 152-161. https://doi.org/10.1016/j.compedu.2014.08.019

Donoghue, G. M., & Hattie, J. A. C. (2021). A meta-analysis of ten learning techniques. Frontiers in Education, 6, 581216. https://doi.org/10.3389/feduc.2021.581216

Hays, M. J., Kornell, N., & Bjork, R. A. (2013). When and why a failed test potentiates the effectiveness of subsequent study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 39(1), 290-296. https://doi.org/10.1037/a0028468

Kang, S. H. K., McDermott, K. B., & Roediger, H. L., III. (2007). Test format and corrective feedback modify the effect of testing on long-term retention. European Journal of Cognitive Psychology, 19(4-5), 528-558. https://doi.org/10.1080/09541440601056620

Karpicke, J. D., & Aue, W. R. (2015). The testing effect is alive and well with complex materials. Educational Psychology Review, 27(2), 317-326. https://doi.org/10.1007/s10648-015-9309-3

Karpicke, J. D., & Blunt, J. R. (2011). Retrieval practice produces more learning than elaborative studying with concept mapping. Science, 331(6018), 772-775. https://doi.org/10.1126/science.1199327

Karpicke, J. D., Butler, A. C., & Roediger, H. L., III. (2009). Metacognitive strategies in student learning: Do students practise retrieval when they study on their own? Memory, 17(4), 471-479. https://doi.org/10.1080/09658210802647009

Karpicke, J. D., & Roediger, H. L., III. (2007). Expanding retrieval practice promotes short-term retention, but equally spaced retrieval enhances long-term retention. Journal of Experimental Psychology: Learning, Memory, and Cognition, 33(4), 704-719. https://doi.org/10.1037/0278-7393.33.4.704

Karpicke, J. D., & Roediger, H. L., III. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966-968. https://doi.org/10.1126/science.1152408

Kimball, D. R., & Metcalfe, J. (2003). Delaying judgments of learning affects memory, not metamemory. Memory & Cognition, 31(6), 918-929. https://doi.org/10.3758/BF03196445

Koh, A. W. L., Lee, S. C., & Lim, S. W. H. (2018). The learning benefits of teaching: A retrieval practice hypothesis. Applied Cognitive Psychology, 32(3), 401-410. https://doi.org/10.1002/acp.3410

Koriat, A., & Bjork, R. A. (2005). Illusions of competence in monitoring one's knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(2), 187-194. https://doi.org/10.1037/0278-7393.31.2.187

Kornell, N., & Bjork, R. A. (2007). The promise and perils of self-regulated study. Psychonomic Bulletin & Review, 14(2), 219-224. https://doi.org/10.3758/BF03194055

Leroy, S. (2009). Why is it so hard to do my work? The challenge of attention residue when switching between work tasks. Organizational Behavior and Human Decision Processes, 109(2), 168-181. https://doi.org/10.1016/j.obhdp.2009.04.002

Kornell, N., Hays, M. J., & Bjork, R. A. (2009). Unsuccessful retrieval attempts enhance subsequent learning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(4), 989-998. https://doi.org/10.1037/a0015729

Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. European Journal of Social Psychology, 40(6), 998-1009. https://doi.org/10.1002/ejsp.674

Landauer, T. K., & Bjork, R. A. (1978). Optimum rehearsal patterns and name learning. In M. M. Gruneberg et al. (Eds.), Practical aspects of memory (pp. 625-632). Academic Press.

Latimier, A., Peyre, H., & Ramus, F. (2021). A meta-analytic review of the benefit of spacing out retrieval practice episodes on retention. Educational Psychology Review, 33, 959-987. https://doi.org/10.1007/s10648-020-09572-8

Mazeas, A., Duclos, M., Pereira, B., & Chalabaev, A. (2022). Evaluating the effectiveness of gamification on physical activity: Systematic review and meta-analysis of randomized controlled trials. Journal of Medical Internet Research, 24(1), e26779. https://doi.org/10.2196/26779

Mekler, E. D., Brühlmann, F., Tuch, A. N., & Opwis, K. (2017). Towards understanding the effects of individual gamification elements on intrinsic motivation and performance. Computers in Human Behavior, 71, 525-534. https://doi.org/10.1016/j.chb.2015.08.048

Morris, C. D., Bransford, J. D., & Franks, J. J. (1977). Levels of processing versus transfer appropriate processing. Journal of Verbal Learning and Verbal Behavior, 16(5), 519-533. https://doi.org/10.1016/S0022-5371(77)80016-9

Nelson, T. O., & Dunlosky, J. (1991). When people's judgments of learning (JOLs) are extremely accurate at predicting subsequent recall: The "delayed-JOL effect." Psychological Science, 2(4), 267-270. https://doi.org/10.1111/j.1467-9280.1991.tb00147.x

Nestojko, J. F., Bui, D. C., Kornell, N., & Bjork, E. L. (2014). Expecting to teach enhances learning and organization of knowledge in free recall of text passages. Memory & Cognition, 42(7), 1038-1048. https://doi.org/10.3758/s13421-014-0416-z

Ophir, E., Nass, C., & Wagner, A. D. (2009). Cognitive control in media multitaskers. Proceedings of the National Academy of Sciences, 106(37), 15583-15587. https://doi.org/10.1073/pnas.0903620106

Philpott, A., & Son, J.-B. (2022). Leaderboards in an EFL course: Student performance and motivation. Computers & Education, 190, 104605. https://doi.org/10.1016/j.compedu.2022.104605

Rawson, K. A., & Dunlosky, J. (2011). Optimizing schedules of retrieval practice for durable and efficient learning: How much is enough? Journal of Experimental Psychology: General, 140(3), 283-302. https://doi.org/10.1037/a0022547

Rawson, K. A., Dunlosky, J., & Sciartelli, S. M. (2013). The power of successive relearning: Improving performance on course exams and long-term retention. Educational Psychology Review, 25(4), 523-548. https://doi.org/10.1007/s10648-013-9240-4

Richland, L. E., Kornell, N., & Kao, L. S. (2009). The pretesting effect: Do unsuccessful retrieval attempts enhance learning? Journal of Experimental Psychology: Applied, 15(3), 243-257. https://doi.org/10.1037/a0016496

Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255. https://doi.org/10.1111/j.1467-9280.2006.01693.x

Roediger, H. L., III, & Marsh, E. J. (2005). The positive and negative consequences of multiple-choice testing. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(5), 1155-1159. https://doi.org/10.1037/0278-7393.31.5.1155

Rohrer, D., Dedrick, R. F., & Stershic, S. (2015). Interleaved practice improves mathematics learning. Journal of Educational Psychology, 107(3), 900-908. https://doi.org/10.1037/edu0000001

Rohrer, D., & Taylor, K. (2006). The effects of overlearning and distributed practice on the retention of mathematics knowledge. Applied Cognitive Psychology, 20(9), 1209-1224. https://doi.org/10.1002/acp.1266

Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432-1463. https://doi.org/10.1037/a0037559

Taylor, K., & Rohrer, D. (2010). The effects of interleaved practice. Applied Cognitive Psychology, 24(6), 837-848. https://doi.org/10.1002/acp.1598

Sailer, M., & Homner, L. (2020). The gamification of learning: A meta-analysis. Educational Psychology Review, 32(1), 77-112. https://doi.org/10.1007/s10648-019-09498-w

Silverman, J., & Barasch, A. (2023). On or off track: How (broken) streaks affect consumer decisions. Journal of Consumer Research, 49(6), 1095-1117. https://doi.org/10.1093/jcr/ucac029

Slamecka, N. J., & Graf, P. (1978). The generation effect: Delineation of a phenomenon. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592-604. https://doi.org/10.1037/0278-7393.4.6.592

Smith, M. A., Roediger, H. L., III, & Karpicke, J. D. (2013). Covert retrieval practice benefits retention as much as overt retrieval practice. Journal of Experimental Psychology: Learning, Memory, and Cognition, 39(6), 1712-1725. https://doi.org/10.1037/a0033569

Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176-199. https://doi.org/10.1177/1745691615569000

Thiede, K. W., Anderson, M. C. M., & Therriault, D. J. (2003). Accuracy of metacognitive monitoring affects learning of texts. Journal of Educational Psychology, 95(1), 66-73. https://doi.org/10.1037/0022-0663.95.1.66

Tran, R., Rohrer, D., & Pashler, H. (2015). Retrieval practice: The lack of transfer to deductive inferences. Psychonomic Bulletin & Review, 22(1), 135-140. https://doi.org/10.3758/s13423-014-0646-x

van Gog, T., & Sweller, J. (2015). Not new, but nearly forgotten: The testing effect decreases or even disappears as the complexity of learning materials increases. Educational Psychology Review, 27(2), 247-264. https://doi.org/10.1007/s10648-015-9310-x

Ye, J., Su, J., & Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4381-4390. https://doi.org/10.1145/3534678.3539081