Cultivar: the evidence

Cultivating You. A habit tracker where the day is a plant you feed and the year is a garden you can look back over.

Cultivar is built by one person, a licensed master social worker (LMSW) with three undergraduate and two graduate degrees. Every rule in the app comes from the research on this page: what a missed day costs, why there is no streak to lose, and why it asks how a habit felt and when it happened. Each finding carries a confidence tag, and the places where the evidence cuts against the app are marked rather than left out.

Cultivar is a habit tracker, not therapy and not a clinical tool. Nothing in the app or on this page is medical or mental-health advice.

How to read the confidence tags

Tag Means
Strong Primary source read, or a meta-analysis over many trials
Moderate One or two decent studies, or a well-replicated effect from an adjacent field
Weak Industry data, secondary reporting, or a framework with little direct evidence

Nothing here is tagged Strong on the strength of a summary alone.


1. Habit formation is an asymptotic curve, not a countdown

Strong. Lally et al. (2010) is the study everything else cites, and it is worth reading past the headline. 96 volunteers chose a daily eating, drinking or activity behaviour tied to a cue ("after breakfast"), and rated automaticity daily for 84 days.

For Cultivar. The Garden is already the right shape for this — it shows accumulation without implying a finish line. What the app should never do is imply a countdown: no "21 days to go", no progress bar toward habit-hood. Somebody at day 20 with a patchy record is on an ordinary curve, and telling them so is both true and kind.

2. One missed day costs almost nothing. A missed week is different.

Strong, and this is the single most important finding in the file.

Lally identified 140 missed opportunities across 55 participants and measured automaticity either side:

The authors' conclusion: "missing one opportunity did not materially affect the habit formation process."

But the same paper draws the contrast. Armitage (2005) assessed behaviour in weekly blocks and defined a lapse as not attending for a week — and there, lapses were negative predictors of future performance. Lally's reading: "missing one opportunity does not preclude habit formation, but missing a week's worth of opportunities reduces the likelihood of future performance."

For Cultivar. This is empirical backing for the whole harm-reduction stance, and it is sharper than the stance currently is. A single miss is genuinely noise —the app should not react to it at all. A gap approaching a week is the first point at which the evidence says something real is happening. The app's "wilting" signal, which fires at three missed due days, sits sensibly between those, and is defensible as an early, quiet signal rather than an alarm.

It also means the app's current behaviour is slightly too eager in one place: a single bare day still darkens the grid cell, which is honest as a record, but nothing in the reflection should ever comment on one.

3. Context stability beats willpower, and time is part of context

Moderate→Strong. Wood & Neal's model is that habits are cue-contingent: behaviour gets triggered by a stable context, and strong habits run largely independently of current goals. The practical consequence is that habits are built by repetition in a stable context, and broken by context change — the habit-discontinuity findings (moving house, changing university, lockdown) are the clean demonstration.

A 2022 study tested this directly (Frontiers in Psychology; Study 1: 95 participants, 2,482 repetitions, contexts experimentally stabilised or varied; Study 2: 218 app users, 308 habits, 2,368 repetitions). Context stability predicted both higher automaticity and better goal attainment, with automaticity partially mediating. Context here explicitly includes time of day, not just place.

On timing specifically (Moderate): consistency of hour matters more than which hour. Exercising at a consistent time predicts greater automaticity and higher activity levels regardless of when that time is; Fournier et al. (2017, N = 48) found morning stretching reached habit status faster than evening (106 vs 154 days, extrapolated — see §10), but the consistency effect is the robust one.

For Cultivar. This is the strongest argument yet for doing something with completedAt, which the app already stores and ignores. The derived insight is not "most of your ticks land before nine" as a curiosity — it is that a habit performed at a scattered hour is predicted to take longer to become automatic than the same habit performed at a consistent one, and that is genuinely actionable without being coercive. It also reframes a habit that has gone quiet after a life change as a context problem rather than a motivation problem, which is both truer and kinder.

4. If-then plans work — but they are not a shortcut to automaticity

Strong on the effect, Moderate on the mechanism, and the mechanism is the interesting part.

Implementation intentions ("if situation X, then I will do Y") are among the best-evidenced techniques in behaviour change: across 642 independent tests, effects run d = .27 to .66, with roughly d = 0.59 for health behaviours. Effects are larger when the plan is genuinely contingent, the person is motivated, and the plan is rehearsed.

van Timmeren & de Wit (2023) tested whether they create instant habits — stimulus–response links insensitive to outcome value — in a single 70-minute laboratory session (two experiments, N = 60 and N = 30). If-then plans gave a transient accuracy edge during training and worse accuracy at test, but the deficit was general, hitting trials where the planned response was still right, and the predicted rigidity interaction was absent (Bayes factors 1.72 and 1.04). Their reading: the plans "failed to instal 'instant habits'"; the harm ran through people learning less about outcomes because the plan had done the learning for them. Their caution is scoped to agents who "do not already possess perfect knowledge of behavioural contingencies" — sport, aviation, surgery.

For Cultivar. Cultivar's schedule (this habit, these weekdays) is a weak implementation intention already, and the effect on goal attainment is real. An earlier version of this paragraph used van Timmeren & de Wit to argue that cue fields and "after X I will Y" recipes would trade away the app's flexibility; a close read (§10) shows the paper says nothing of the kind about a self-chosen daily behaviour, and the argument is withdrawn. The case against an anchor field now rests on a different finding — that a routine cue is no better than any other consistent cue — and §10 makes it.

5. Rigid routines break; flexible ones survive disruption

Weak-to-Moderate, and flagged as such because the most-quoted result here is secondary reporting rather than a paper I could read. The claim in circulation is that people given a flexible plan were more than twice as likely to still be exercising after four weeks than those given a rigid routine, on the reasoning that a rigid routine has more surface area for disruption to hit.

The adjacent, better-evidenced finding is that psychological rigidity predicts worse outcomes across a range of processes, and that flexible coping beats rigid adherence to any single strategy.

For Cultivar. Treat as suggestive, not settled. It points the same way as everything else here — rest days that carry, schedules that can shrink, a plant that resets rather than accumulates — so it reinforces rather than decides. Do not cite the 2× figure in the app or in copy.

6. Tracking itself changes behaviour — and the effect decays

Moderate. Self-monitoring is defined in Michie's taxonomy as keeping a record of a behaviour in order to change it, and it is one of the stronger single predictors of behaviour change. The important qualifications:

For Cultivar. Two consequences. First, the tracker is not neutral furniture — the act of recording is itself part of the intervention, which raises the stakes on making recording cheap. The widget is the app's best behaviour-change feature and should be defended as such. Second, decay is the expected failure mode, so the thing to watch is not whether a habit's rate drops but whether tracking stops.

7. Streaks: the mechanism is real, and so is the harm

Moderate on the harm, Weak-but-instructive on the industry data.

The theoretical case against is the overjustification effect: introducing an extrinsic reward for an intrinsically motivated behaviour reduces intrinsic motivation. Applied to habit apps, streaks and punishments recruit introjected regulation — guilt and shame — rather than identified or intrinsic motivation, and produce compliance rather than sustained change. Gamification can support or thwart autonomy and competence depending entirely on execution.

The industry data cuts an interesting way. Duolingo's streak freeze reduced churn by 21% among users at risk of breaking a streak, and users offered it were more likely to return. The reported finding is that by permitting breaks, users did more learning in the long run.

Read that carefully: the most successful streak mechanic in the industry is the one that lets the streak survive a missed day. The value was in removing the cliff, not in the counter.

For Cultivar. The app already has the good half of this — rest days carry the streak, there is no red X, and the streak now sits below the rate on Patterns. The evidence supports going further rather than turning back: nothing in the literature argues that a visible count adds anything once the cliff is gone. If the streak ever starts driving behaviour visibly, that is a reason to demote it again, not to celebrate.

8. Complexity sets the pace, and consistency sets the ceiling

Moderate. In Lally, median time to plateau by behaviour type: drinking 59 days, eating 65, exercise 91. Not statistically significant with that sample, but in the predicted direction, and consistent with Wood & Neal's argument that complex behaviours develop goal-directed automaticity rather than habit proper.

Separately: four participants fit the model well but plateaued low, and they had performed the behaviour significantly less often. Lally suggests a threshold of performance below which maximum habit strength is curtailed — above it, something other than raw repetition count determines the ceiling.

For Cultivar. Two usable facts. A demanding habit still feeling effortful at two months is normal and expected, and saying so is one of the few genuinely reassuring things the app can say that is also true. And where a habit is being kept inconsistently, the lever is consistency rather than intensity — which is exactly the argument for suggesting a smaller ask, now with evidence behind it rather than taste alone.

9. Demographics would not move the needle. Two non-demographic things would.

Moderate, and the shape of the evidence matters more than any single number.

The question: would knowing the user's age, sex, occupation, income or anything similar help them keep a habit? The honest answer is no, and for a more specific reason than "no evidence".

The determinant list does not contain a single demographic. The most recent systematic review and meta-analysis (Singh et al. 2024; 20 studies, 2,601 participants, mean ages 21.5–73.5) identifies the things that significantly influence habit strength: frequency, timing, type of habit, individual choice, affective judgement, behavioural regulation, and preparatory habits. Age, sex, education, income and BMI were not examined as moderators at all. That is absence of evidence rather than evidence of absence — but across this literature nobody treats demographics as the lever, and the variables that keep reappearing are all things about the behaviour and its context rather than about the person.

Lally points the same way. The spread in time-to-plateau was 18 to 254 days, and what correlated with it was behaviour complexity and consistency of performance, not who the person was.

The tailoring literature makes the distinction explicit. Computer-tailored interventions do beat generic ones, but the meta-analytic finding is that dynamic tailoring — adapting to the person's own behavioural pattern over time — outperforms static tailoring on demographics. Static tailoring variables are exactly the age-and- gender kind; dynamic ones are what the person actually did last week.

And for a single-user app the argument is stronger still. Population moderators exist to allocate interventions across groups. Cultivar has n = 1 and direct observation of the behaviour, which dominates any demographic proxy. Every demographic worth having is a weak stand-in for something the app already measures:

Demographic What it is a proxy for What Cultivar already has
Age Morning preference, routine stability completedAt median and spread
Occupation / shift pattern Which weekdays are hard the habit × weekday matrix
Sex Nothing reliable for habit formation —
Baseline fitness / BMI How demanding a habit is target, unit, note, and the rate

Chronotype is the closest thing to a real exception, and it does not survive contact either. The evidence for matching exercise to chronotype is mostly about health outcomes — blood pressure, glucose, cardiovascular risk — rather than adherence, and the adherence-flavoured recommendation that comes out of it is "let people choose when they feel best", which is autonomy, not data collection. A questionnaire would ask someone to self-report a preference the app can already see in their actual timestamps.

There is also a cost, and it is not only onboarding friction. Advice conditioned on a demographic is advice conditioned on a stereotype — "people your age often find…" is precisely the coach's voice this app exists to avoid, and it would be less accurate than the person's own record.

The two things worth collecting instead, neither of them demographic

1. Affective judgement — how the behaviour felt. This is the strongest candidate by some distance. Singh et al. name affective judgement (enjoyment of the behaviour) as a significant mediator of physical activity habit formation, alongside behavioural regulation and preparatory habit. A habit kept at 90% that feels like a chore and one kept at 90% that is enjoyed have different futures. Cultivar now asks how each habit felt: one optional tap after it is done, never required, never scored. The Mood plot on Patterns crosses that with how often the habit is kept, which is what tells the two apart.

2. Context — what the habit hangs off. Context stability is the best-evidenced accelerator in the file (§3), and time is only one dimension of it. The app sees when but not where or after what. Capture it as an observation, not as an if-then contract — not for the reason §4 used to give, which §10 withdraws, but because the routine-versus-time trial found the two interchangeable and the one anchoring trial found self-chosen anchors did nothing. A field asking "after what?" would be adopting the recipe on Fogg's authority.

Habit-specific self-efficacy is a real third candidate — it predicts automaticity and the relationship is reciprocal (b = 0.416 and 0.327 in a 196-user, 2,132-repetition study), forming a self-amplifying loop. But it is also the one most safely left unmeasured: asking "how confident are you?" invites the self-punishing reader this app is written for to score themselves, and the loop is fed by mastery experiences the app already generates simply by recording that something got done.

Recommendation: collect no demographics at all. If exactly one field is ever added, make it affective judgement, optional and one tap.

10. Anchor habits and habit stacking: the cue is the finding, the routine is a convenience

Strong that a consistent cue builds automaticity. Weak that a routine cue beats any other consistent cue. Weak on the Tiny Habits method itself.

Fogg's recipe is "After I [existing routine], I will [tiny behaviour]"; Clear's "habit stacking" is the same recipe with Fogg's credit attached. The academic version is cue-dependent repetition with an event-based rather than a time-based cue, and the literature separates the two claims cleanly.

The cue is where the evidence is.

The routine is not where the evidence is.

Anchors are implementation intentions, and §4 has to be corrected. An anchor is an if-then plan whose "if" is the end of an existing act; Keller's routine arm is the recipe and its time arm is the conventional plan. The implementation-intention meta-analyses (Adriaanse et al. 2011, 23 studies, d = 0.43 for eating; Carrero et al. 2019, k = 70, d = 0.33) never model cue type, so they support "plan a cue", not "the cue should be a routine". And the paper this file used against anchors does not carry that weight. van Timmeren & de Wit ran a single 70-minute laboratory session in which an if-then plan was handed to people as a substitute for learning which of eight ice-cream vans signalled which outcome. The test deficit was general — it hit trials where the planned response was still right — and the predicted rigidity interaction was absent (Bayes factors 1.72 and 1.04, inconclusive both ways). The authors' own conclusion is that the plans "failed to instal 'instant habits'", that the harm ran through reduced learning of outcomes, and that caution applies "to situations where the agent does not already possess perfect knowledge of behavioural contingencies" — their examples are sport, aviation and surgery. A person who chooses to drink water after breakfast has nothing left to learn about the contingency. The paper says nothing about anchors in daily life, and the earlier reading of it is withdrawn.

For Cultivar. Three things follow, and the first is the important one.

11. Keystone habits: the cascade has lost its mechanism, and the spillovers that survive are small

Weak on Duhigg's claim and on the mechanism he gave it. Moderate that small, same-direction spillovers between health behaviours exist and fade. Moderate on preparatory habits, which are the one cascade with a decent effect size — and they are a chain within one behaviour, not across behaviours.

Duhigg's chapter says some habits — exercise, family dinners, making the bed, Alcoa's safety programme — "start a chain reaction" through small wins and strengthened willpower. He hedges the causation in one sentence and asserts it in the next. Habit science has not adopted the term: Gardner et al.'s 2023 "twenty-one questions" paper never uses it, and the closest it comes is an open research question about whether a "higher-order habit" promotes more specific ones. Of 59 works in OpenAlex using the phrase, none is by Wood, Gardner, Rebar, Verplanken, Lally or Rhodes.

What Duhigg cites, and what happened to it.

The real phenomenon is behavioural spillover, and it is modest.

Negative spillover is real, conditional, and smaller than feared.

The one cascade with an effect size is local. A preparatory act — packing the bag, laying out the mat — cues the target behaviour: preparatory habit predicted exercise at β = .20 (N = 181) where execution habit did not, and teaching cue use and consistency to new gym members raised accelerometer-measured activity by d = 0.39 at eight weeks (N = 94). Singh et al. list preparatory habits among the determinants of habit strength (§9). This is the keystone idea shrunk to its evidenced size: a small act that reliably starts a larger one, within the same behaviour.

Identity is the plausible route across behaviours, and it nets to nothing on its own. Habit and identity correlate at r = 0.55 across 19 studies and 13,340 people — Strong as an association, mostly cross-sectional. But reminding people of a past good behaviour raises identity and lowers guilt, and the two cancel: net spillover ≈ 0 (Lacasse 2016, N = 377 and 172; replicated 2021). Only an explicit label ("environmentalist") kept the identity gain without the guilt drop. A field experiment adopting one new behaviour for three weeks (N = 125) found spillover "limited and relatively small".

For Cultivar.


What this means the reflection should look for

The reflection has been producing restatements because the summary gives it aggregates and the prompt asks for "an observation". The research above names the observations that would actually be worth making. This is the catalogue to build toward.

# Pattern Grounded in Data needed Status
1 A day asks for more than it is getting §8 consistency ceiling weekday load + per-habit rate on that day built
2 A habit is kept at a scattered hour vs a consistent one §3 context stability completedAt fitted per habit (domain/hours: von Mises mixture on the last 60 days, change points over everything); two settled hours read as a routine, a moved hour reads as a move with its before and after built
3 A gap is approaching a week, not just a day §2 Lally vs Armitage consecutive missed due days partly — dated gap spans are sent; the wilting signal still fires at 3 days
4 A demanding habit is slower to settle, and that is normal §8 complexity habit age + target/unit as a proxy for effort built (age, target, note all sent)
5 Two habits rise and fall together §3 shared context per-day co-occurrence between roots built — pairs ranked by distance from independence
6 A habit went quiet when something else changed §3 habit discontinuity change-point against schedule edits partly — dated gaps are sent, schedule edits are not
7 Early repetitions are buying more than later ones §1 asymptotic curve habit age built
8 One habit is kept more on days another came first §10 instigation, §11 preparatory chain completedAt order within a pair not built — Together uses co-occurrence, not order

Verified live against Claude Opus 5 once the data was there. Before, the year window said "Reading is the only thing due on Saturday, and that day sits at 50 percent" — a figure read back. After: "Stretch reads as the habit kept least often, but on Tuesday and Saturday it runs at 60 percent, the same rate meditation holds on six days of the week; the 40 percent comes from Thursday alone." That is an aggregate being corrected by its own breakdown, which is the thing the five rule-based cards structurally cannot do.

The distinction the app keeps failing to draw: an observation restates a number in prose; a derived insight proposes a mechanism. "Thursday sits at 50%" is the former. "Thursday is the only day carrying four habits, and the two that fail there are the two that need the most setup" is the latter. Only the second gives the person anything to act on, and only the second earns the cost of the call.

What the research does not support

Worth writing down, because these are the things that will keep being suggested:

Open questions

Things this app is in an unusually good position to answer for its own user, and currently cannot:


Sources

More

Privacy and support