Cultivating You. A habit tracker where the day is a plant you feed and the year is a garden you can look back over.
Cultivar is built by one person, a licensed master social worker (LMSW) with three undergraduate and two graduate degrees. Every rule in the app comes from the research on this page: what a missed day costs, why there is no streak to lose, and why it asks how a habit felt and when it happened. Each finding carries a confidence tag, and the places where the evidence cuts against the app are marked rather than left out.
Cultivar is a habit tracker, not therapy and not a clinical tool. Nothing in the app or on this page is medical or mental-health advice.
How to read the confidence
tags
Tag
Means
Strong
Primary source read, or a meta-analysis over many trials
Moderate
One or two decent studies, or a well-replicated effect from an
adjacent field
Weak
Industry data, secondary reporting, or a framework with little
direct evidence
Nothing here is tagged Strong on the strength of a summary alone.
1.
Habit formation is an asymptotic curve, not a countdown
Strong. Lally et al. (2010) is the study everything
else cites, and it is worth reading past the headline. 96 volunteers
chose a daily eating, drinking or activity behaviour tied to a cue
("after breakfast"), and rated automaticity daily for 84 days.
The median time to reach 95% of a person's own plateau was
66 days, with a range of 18 to 254 days and
quartiles of 39–102. The famous "66 days" is a median of a wildly
dispersed distribution, not a target.
Automaticity follows Mitscherlich's law of diminishing
returns — early repetitions buy far more than later ones, and
there is a ceiling past which more repetition adds nothing. The
nonlinear model beat a linear one decisively (Z = −5.765,
p < .001).
The curve fit well for only 39 of 82 participants
(48%). Roughly half the people in a study specifically
designed to build a habit never performed it consistently enough to
form one. This is the normal case, not the failure case.
For Cultivar. The Garden is already the right shape
for this — it shows accumulation without implying a finish line. What
the app should never do is imply a countdown: no "21 days to go", no
progress bar toward habit-hood. Somebody at day 20 with a patchy record
is on an ordinary curve, and telling them so is both true and kind.
2.
One missed day costs almost nothing. A missed week is different.
Strong, and this is the single most important
finding in the file.
Lally identified 140 missed opportunities across 55
participants and measured automaticity either side:
Immediately after a miss, automaticity fell by 0.29
points on a 0–42 scale.
Where the behaviour resumed the next day (N = 67), the
before-to-after change was an increase of 0.55 points —
not significant on a Wilcoxon signed rank test.
Over three consecutive performed days the average increase
was 0.79 points. So a miss costs roughly a quarter of a
point of progress against a day that went well.
The timing of the miss did not matter (r = 0.099,
p = 0.246): missing early is no worse than missing late.
The authors' conclusion: "missing one opportunity did not
materially affect the habit formation process."
But the same paper draws the contrast. Armitage
(2005) assessed behaviour in weekly blocks and defined a lapse as
not attending for a week — and there, lapses
were negative predictors of future performance. Lally's
reading: "missing one opportunity does not preclude habit formation, but
missing a week's worth of opportunities reduces the likelihood of future
performance."
For Cultivar. This is empirical backing for the
whole harm-reduction stance, and it is sharper than the stance currently
is. A single miss is genuinely noise —the app should not react to it at
all. A gap approaching a week is the first point at which the evidence
says something real is happening. The app's "wilting" signal, which
fires at three missed due days, sits sensibly between those,
and is defensible as an early, quiet signal rather than an alarm.
It also means the app's current behaviour is slightly too eager in
one place: a single bare day still darkens the grid cell, which is
honest as a record, but nothing in the reflection should ever comment on
one.
3.
Context stability beats willpower, and time is part of context
Moderate→Strong. Wood & Neal's model is that
habits are cue-contingent: behaviour gets triggered by a stable context,
and strong habits run largely independently of current goals. The
practical consequence is that habits are built by repetition in a
stable context, and broken by context change — the
habit-discontinuity findings (moving house, changing university,
lockdown) are the clean demonstration.
A 2022 study tested this directly (Frontiers in Psychology; Study 1:
95 participants, 2,482 repetitions, contexts experimentally stabilised
or varied; Study 2: 218 app users, 308 habits, 2,368 repetitions).
Context stability predicted both higher automaticity and better
goal attainment, with automaticity partially mediating. Context
here explicitly includes time of day, not just
place.
On timing specifically (Moderate): consistency of
hour matters more than which hour. Exercising at a consistent time
predicts greater automaticity and higher activity levels regardless of
when that time is; Fournier et al. (2017, N = 48) found morning
stretching reached habit status faster than evening (106 vs 154 days,
extrapolated — see §10), but the consistency effect is the robust
one.
For Cultivar. This is the strongest argument yet for
doing something with completedAt, which the app already
stores and ignores. The derived insight is not "most of your ticks land
before nine" as a curiosity — it is that a habit performed at a
scattered hour is predicted to take longer to become automatic than the
same habit performed at a consistent one, and that is genuinely
actionable without being coercive. It also reframes a habit that has
gone quiet after a life change as a context problem rather than a
motivation problem, which is both truer and kinder.
4.
If-then plans work — but they are not a shortcut to automaticity
Strong on the effect, Moderate on the mechanism, and the
mechanism is the interesting part.
Implementation intentions ("if situation X, then I will do Y") are
among the best-evidenced techniques in behaviour change: across
642 independent tests, effects run d = .27 to
.66, with roughly d = 0.59 for health
behaviours. Effects are larger when the plan is genuinely contingent,
the person is motivated, and the plan is rehearsed.
van Timmeren & de Wit (2023) tested whether they create
instant habits — stimulus–response links insensitive to outcome
value — in a single 70-minute laboratory session (two experiments,
N = 60 and N = 30). If-then plans gave a transient
accuracy edge during training and worse accuracy at
test, but the deficit was general, hitting trials where the
planned response was still right, and the predicted rigidity interaction
was absent (Bayes factors 1.72 and 1.04). Their reading: the plans
"failed to instal 'instant habits'"; the harm ran through people
learning less about outcomes because the plan had done the learning for
them. Their caution is scoped to agents who "do not already possess
perfect knowledge of behavioural contingencies" — sport, aviation,
surgery.
For Cultivar. Cultivar's schedule (this habit, these
weekdays) is a weak implementation intention already, and the effect on
goal attainment is real. An earlier version of this paragraph used van
Timmeren & de Wit to argue that cue fields and "after X I will Y"
recipes would trade away the app's flexibility; a close read (§10) shows
the paper says nothing of the kind about a self-chosen daily behaviour,
and the argument is withdrawn. The case against an anchor field now
rests on a different finding — that a routine cue is no better than any
other consistent cue — and §10 makes it.
Weak-to-Moderate, and flagged as such because the
most-quoted result here is secondary reporting rather than a paper I
could read. The claim in circulation is that people given a flexible
plan were more than twice as likely to still be exercising after four
weeks than those given a rigid routine, on the reasoning that a rigid
routine has more surface area for disruption to hit.
The adjacent, better-evidenced finding is that psychological rigidity
predicts worse outcomes across a range of processes, and that flexible
coping beats rigid adherence to any single strategy.
For Cultivar. Treat as suggestive, not settled. It
points the same way as everything else here — rest days that carry,
schedules that can shrink, a plant that resets rather than accumulates —
so it reinforces rather than decides. Do not cite the 2× figure in the
app or in copy.
6.
Tracking itself changes behaviour — and the effect decays
Moderate. Self-monitoring is defined in Michie's
taxonomy as keeping a record of a behaviour in order to change it, and
it is one of the stronger single predictors of behaviour change. The
important qualifications:
It is much more effective combined with other
self-regulation techniques from control theory than alone. One
review found no effect for self-monitoring in isolation.
Engagement decays. The consistent barriers are time
demands, perceived burden, and waning novelty.
For Cultivar. Two consequences. First, the tracker
is not neutral furniture — the act of recording is itself part of the
intervention, which raises the stakes on making recording cheap. The
widget is the app's best behaviour-change feature and should be defended
as such. Second, decay is the expected failure mode, so the thing to
watch is not whether a habit's rate drops but whether tracking
stops.
7. Streaks:
the mechanism is real, and so is the harm
Moderate on the harm, Weak-but-instructive on the industry
data.
The theoretical case against is the overjustification effect:
introducing an extrinsic reward for an intrinsically motivated behaviour
reduces intrinsic motivation. Applied to habit apps, streaks and
punishments recruit introjected regulation — guilt and
shame — rather than identified or intrinsic motivation, and produce
compliance rather than sustained change. Gamification can support or
thwart autonomy and competence depending entirely on execution.
The industry data cuts an interesting way. Duolingo's streak
freeze reduced churn by 21% among users at risk of breaking a
streak, and users offered it were more likely to return. The
reported finding is that by permitting breaks, users did more
learning in the long run.
Read that carefully: the most successful streak mechanic in the
industry is the one that lets the streak survive a missed day.
The value was in removing the cliff, not in the counter.
For Cultivar. The app already has the good half of
this — rest days carry the streak, there is no red X, and the streak now
sits below the rate on Patterns. The evidence supports going further
rather than turning back: nothing in the literature argues that a
visible count adds anything once the cliff is gone. If the
streak ever starts driving behaviour visibly, that is a reason to demote
it again, not to celebrate.
8.
Complexity sets the pace, and consistency sets the ceiling
Moderate. In Lally, median time to plateau by
behaviour type: drinking 59 days, eating 65, exercise
91. Not statistically significant with that sample, but in the
predicted direction, and consistent with Wood & Neal's argument that
complex behaviours develop goal-directed automaticity rather than habit
proper.
Separately: four participants fit the model well but plateaued
low, and they had performed the behaviour significantly less
often. Lally suggests a threshold of performance below which
maximum habit strength is curtailed — above it, something other
than raw repetition count determines the ceiling.
For Cultivar. Two usable facts. A demanding habit
still feeling effortful at two months is normal and
expected, and saying so is one of the few genuinely reassuring
things the app can say that is also true. And where a habit is being
kept inconsistently, the lever is consistency rather than intensity —
which is exactly the argument for suggesting a smaller ask, now
with evidence behind it rather than taste alone.
9.
Demographics would not move the needle. Two non-demographic things
would.
Moderate, and the shape of the evidence matters more than any
single number.
The question: would knowing the user's age, sex, occupation, income
or anything similar help them keep a habit? The honest answer is no, and
for a more specific reason than "no evidence".
The determinant list does not contain a single
demographic. The most recent systematic review and
meta-analysis (Singh et al. 2024; 20 studies, 2,601 participants, mean
ages 21.5–73.5) identifies the things that significantly influence habit
strength: frequency, timing, type of habit, individual choice,
affective judgement, behavioural regulation, and preparatory
habits. Age, sex, education, income and BMI were not examined
as moderators at all. That is absence of evidence rather than evidence
of absence — but across this literature nobody treats demographics as
the lever, and the variables that keep reappearing are all things about
the behaviour and its context rather than about the person.
Lally points the same way. The spread in time-to-plateau was 18 to
254 days, and what correlated with it was behaviour complexity and
consistency of performance, not who the person was.
The tailoring literature makes the distinction
explicit. Computer-tailored interventions do beat generic ones,
but the meta-analytic finding is that dynamic tailoring —
adapting to the person's own behavioural pattern over time — outperforms
static tailoring on demographics. Static tailoring variables
are exactly the age-and- gender kind; dynamic ones are what the person
actually did last week.
And for a single-user app the argument is stronger
still. Population moderators exist to allocate interventions
across groups. Cultivar has n = 1 and direct observation of the
behaviour, which dominates any demographic proxy. Every demographic
worth having is a weak stand-in for something the app already
measures:
Demographic
What it is a proxy for
What Cultivar already has
Age
Morning preference, routine stability
completedAt median and spread
Occupation / shift pattern
Which weekdays are hard
the habit × weekday matrix
Sex
Nothing reliable for habit formation
—
Baseline fitness / BMI
How demanding a habit is
target, unit, note, and the rate
Chronotype is the closest thing to a real exception, and it does not
survive contact either. The evidence for matching exercise to chronotype
is mostly about health outcomes — blood pressure, glucose,
cardiovascular risk — rather than adherence, and the adherence-flavoured
recommendation that comes out of it is "let people choose when they feel
best", which is autonomy, not data collection. A questionnaire would ask
someone to self-report a preference the app can already see in their
actual timestamps.
There is also a cost, and it is not only onboarding
friction. Advice conditioned on a demographic is advice
conditioned on a stereotype — "people your age often find…" is precisely
the coach's voice this app exists to avoid, and it would be
less accurate than the person's own record.
The
two things worth collecting instead, neither of them demographic
1. Affective judgement — how the behaviour felt.
This is the strongest candidate by some distance. Singh et al. name
affective judgement (enjoyment of the behaviour) as a significant
mediator of physical activity habit formation, alongside
behavioural regulation and preparatory habit. A habit kept at 90% that
feels like a chore and one kept at 90% that is enjoyed have different
futures. Cultivar now asks how each habit felt: one optional tap after
it is done, never required, never scored. The Mood plot on Patterns
crosses that with how often the habit is kept, which is what tells the
two apart.
2. Context — what the habit hangs off. Context
stability is the best-evidenced accelerator in the file (§3), and time
is only one dimension of it. The app sees when but not
where or after what. Capture it as an observation, not
as an if-then contract — not for the reason §4 used to give, which §10
withdraws, but because the routine-versus-time trial found the two
interchangeable and the one anchoring trial found self-chosen anchors
did nothing. A field asking "after what?" would be adopting the recipe
on Fogg's authority.
Habit-specific self-efficacy is a real third candidate — it predicts
automaticity and the relationship is reciprocal (b = 0.416 and
0.327 in a 196-user, 2,132-repetition study), forming a self-amplifying
loop. But it is also the one most safely left unmeasured: asking "how
confident are you?" invites the self-punishing reader this app is
written for to score themselves, and the loop is fed by mastery
experiences the app already generates simply by recording that something
got done.
Recommendation: collect no demographics at all. If
exactly one field is ever added, make it affective judgement, optional
and one tap.
10.
Anchor habits and habit stacking: the cue is the finding, the routine is
a convenience
Strong that a consistent cue builds automaticity. Weak that a
routine cue beats any other consistent cue. Weak on the Tiny Habits
method itself.
Fogg's recipe is "After I [existing routine], I will [tiny
behaviour]"; Clear's "habit stacking" is the same recipe with Fogg's
credit attached. The academic version is cue-dependent repetition with
an event-based rather than a time-based cue, and the literature
separates the two claims cleanly.
The cue is where the evidence is.
Lally's protocol required an anchor. The behaviour "had to
be one that … could be performed in response to a salient daily event
(cue) and … had a cue that occurred every day and only once a day" — the
examples are "eating a piece of fruit with lunch" and "running for 15
minutes before dinner". The 66-day median in §1 is already an
anchored-habit number. With no unanchored arm it says nothing about
whether anchoring is faster than anything else.
Cue consistency is the best-replicated predictor of habit strength.
In new gym members (Kaushal & Rhodes 2015, N = 111, 12
weeks) consistency was the largest predictor of the automaticity
trajectory (b = .21, p < .001), ahead of low
complexity and cues, and four or more sessions a week for six weeks was
the threshold — at week 12, 63.8% of the ≥4/week group
were above the habit cut-off against 22.6% below it. A
2026 meta-analysis (14 studies, N = 6,069) puts context
consistency at r = 0.20 with habit strength. Consistency here
means the same time or the same routine; the questionnaire item
names both.
What automates is the cue → start link, not the doing.
Habitual instigation predicts behaviour frequency; habitual execution
does not (Gardner, Phillips & Judah 2016, N = 229; Phillips
& Gardner 2016, N = 123). An anchor targets exactly that
link, which is why the recipe is theoretically sound. Neither study
tests anchoring.
Habit-based packages that bundle "same context every time" with
planning and self-monitoring move automaticity a little and reliably: in
the Ten Top Tips RCT (Beeken et al. 2017, N = 537, 14 GP
practices) automaticity rose +8.45 points over usual
care at three months and weight fell 0.87 kg more (P = 0.004),
with no weight difference left by 24 months; across ten
physical-activity habit interventions (N = 2,349) habit
strength moved SMD 0.31. Anchoring's own share is not
separable from the rest of the bundle.
The routine is not where the evidence is.
The one RCT that isolates the question is null. Keller et al. (2021)
randomised 192 people to plan a new daily nutrition behaviour after a
routine ("after breakfast") or at a clock time ("at 9 am") and followed
135 of them with daily automaticity ratings for 84 days. Month-three
automaticity was 3.74 vs 3.54 on a six-point scale,
habit formation succeeded in 14 vs 13 people (27 of 117
overall — 23%), median time to plateau was 60 vs 59
days, and the range was 4 to 335. The study had over 90% power
for a small difference. What predicted automaticity was enacting the
plan (within-person B = 0.23, p < .001) and
finding the behaviour intrinsically rewarding at baseline (p =
.006). The authors' sentence: "Linking one's nutrition behaviour to a
daily routine or a specific time of day was similarly effective for
habit formation."
The one positive result did not replicate. Judah, Gardner &
Aunger (2013, N = 50) found flossing after brushing
"tended to" form stronger habits than flossing before; the same team's
larger follow-up (2018, N = 118, 16 weeks) found "behaviour
frequency or habit formation did not differ between conditions", and
Gardner, Rebar & Lally's 2022 methods paper calls the 2013
conclusion "less well-founded" because it came from linear models.
Self-chosen anchors did nothing in the only anchoring RCT with a
habit measure. Stecher et al. (2021, 167 randomised, 101 analysed, Calm
app, eight weeks) gave one arm a prescribed anchor ("after I finish in
the bathroom in the morning, I will meditate") and another a
personalised one. The prescribed morning anchor raised daily meditation
(OR 1.14) and automaticity (+4.56 SRBAI vs control);
the personalised-anchor arm did not differ from control on either. Only
30% of anchor participants anchored reliably, almost all in the morning,
and the authors judged eight weeks too short for even those to form a
routine. Their 2024 walking trial (N = 161) found no
habit-strength differences at any point.
The Tiny Habits method has essentially no habit evidence.
The one RCT (Hollingsworth & Redden 2022, N = 154, five
days) measured gratitude and hope — which rose, d = 0.85 — and
never measured automaticity, with attrition of 29–42%. A 2025 scoping
review of Fogg's model found six studies, none using the "After I…"
recipe explicitly and none measuring habit strength. Fogg's own
participant data are unpublished.
Event cues over reminders is plausible and thinly evidenced.
Stawarz, Cox & Blandford (2015) report that reminders supported
repetition but hindered automaticity while event-based cues raised it,
and that of 115 habit apps reviewed none supported event cues — abstract
only, N unknown. Their 2020 study (N = 39, three
weeks) found people rarely chose a routine alone; the ones who kept the
behaviour paired routine with place and object (92% adherence), and only
11 of 39 developed any automaticity.
Morning beats evening within anchored habits. The §3
stretching study is Fournier et al. (2017, N = 48, 90 days):
both arms were anchored (on waking vs before bed), and the 106 vs 154
days are extrapolated beyond the observation window.
Anchors are implementation intentions, and §4 has to be
corrected. An anchor is an if-then plan whose "if" is the end
of an existing act; Keller's routine arm is the recipe and its
time arm is the conventional plan. The implementation-intention
meta-analyses (Adriaanse et al. 2011, 23 studies, d = 0.43 for
eating; Carrero et al. 2019, k = 70, d = 0.33) never
model cue type, so they support "plan a cue", not "the cue should be a
routine". And the paper this file used against anchors does not carry
that weight. van Timmeren & de Wit ran a single 70-minute laboratory
session in which an if-then plan was handed to people as a substitute
for learning which of eight ice-cream vans signalled which outcome. The
test deficit was general — it hit trials where the planned
response was still right — and the predicted rigidity interaction was
absent (Bayes factors 1.72 and 1.04, inconclusive both ways). The
authors' own conclusion is that the plans "failed to instal 'instant
habits'", that the harm ran through reduced learning of outcomes, and
that caution applies "to situations where the agent does not already
possess perfect knowledge of behavioural contingencies" — their examples
are sport, aviation and surgery. A person who chooses to drink water
after breakfast has nothing left to learn about the contingency. The
paper says nothing about anchors in daily life, and the earlier reading
of it is withdrawn.
For Cultivar. Three things follow, and the first is
the important one.
The signal the app already has is the right one. The evidence is
that any consistent cue, enacted daily, builds automaticity,
and that a routine cue is a convenient choice rather than a faster one.
Consistency of completedAt (§3) measures exactly this
without asking anyone anything. It should stay the app's read on the
question.
No anchor field, no recipe builder. Keller says routine and clock
time are interchangeable; Stecher says a self-chosen anchor did nothing
over eight weeks; the Tiny Habits evidence is a gratitude study. A field
that asks "after what?" would be adopting the recipe on Fogg's
authority. §9's "context as an observation" stands, but as an optional
convenience for the person rather than something the evidence asks
for.
The reflection may say the honest version. When a habit is kept at a
scattered hour, "hang it on something that already happens at a fixed
point — a routine or a time, either works" is what Gardner, Lally &
Wardle tell GPs to say and what Keller found. That is the one anchor
claim with evidence behind it, and it is a suggestion about
consistency, not a contract.
11.
Keystone habits: the cascade has lost its mechanism, and the spillovers
that survive are small
Weak on Duhigg's claim and on the mechanism he gave it.
Moderate that small, same-direction spillovers between health behaviours
exist and fade. Moderate on preparatory habits, which are the one
cascade with a decent effect size — and they are a chain within one
behaviour, not across behaviours.
Duhigg's chapter says some habits — exercise, family dinners, making
the bed, Alcoa's safety programme — "start a chain reaction" through
small wins and strengthened willpower. He hedges the causation in one
sentence and asserts it in the next. Habit science has not adopted the
term: Gardner et al.'s 2023 "twenty-one questions" paper never uses it,
and the closest it comes is an open research question about whether a
"higher-order habit" promotes more specific ones. Of 59 works in
OpenAlex using the phrase, none is by Wood, Gardner, Rebar, Verplanken,
Lally or Rhodes.
What Duhigg cites, and what happened to it.
The exercise cascade rests on Oaten & Cheng (2006):
24 sedentary undergraduates in three cohorts of 9, 6
and 9, a wait-list rather than an active control, self-reported smoking,
drinking, spending and study, and an authors' note that "smaller studies
… can show exaggerated effects". Their study-programme paper had
N = 45 and their financial-monitoring paper N = 49.
Miles et al. compute a training effect of d = 8.59 for one of
them, which is not a size psychology produces.
The mechanism was self-control as a trainable muscle, and the muscle
did not survive replication. Ego depletion, which the model presupposes,
came out at d = 0.04 across 23 labs and 2,141 people
(Hagger et al. 2016), d = 0.06 across 36 labs and 3,531
people with Baumeister-approved paradigms (Vohs et al. 2021), and
d = 0.10 in a third (Dang et al. 2021). Training transfer fared
no better: the one large trial with an active control (Miles et al.
2016, N = 174, six weeks) found no effect on any laboratory or
everyday self-control measure, and the two meta-analyses (Friese et al.
2017, 33 studies; Beames et al. 2017, 29 studies) give g =
0.30–0.36 raw, 0.13–0.28 after bias correction, larger
when strength-model proponents were co-authors and larger in published
than unpublished work (0.45 vs 0.17).
Making the bed has no study at all — every figure in circulation is
a self-selected online poll (Hunch.com, ~68,000 users) or a mattress
company's survey. Family dinners largely proxy family quality: in Add
Health (N = 17,977) the benefits shrink to two of nine
interactions at p < .10 once within-family fixed effects are
used. Alcoa is an anecdote. "Small wins" is a 1984 theoretical
essay.
The real phenomenon is behavioural spillover, and it is
modest.
Dolan & Galizzi's taxonomy — promoting (the second
behaviour follows the first), permitting (the first licenses a
lapse), purging (the second undoes the first) — is the useful
frame, and its authors say there is "relatively little systematic
research" on which one occurs when.
Where interventions on one behaviour have been checked for effects
on another, the spillovers are small and transient. A dietary trial in
colorectal cancer patients (n = 469) added 0.18 hours a
day of activity at six months and nothing significant at
twelve; a cluster trial added 1,105 steps a day at three months and none
by six. The Women's Health Initiative dietary arm (n = 20,380)
did not act as a gateway to activity, alcohol or smoking. The first RCT
test of "exercise as a gateway to diet" (N = 280 women, 12
months) found no effect on fruit and vegetables, and more activity went
with more fat intake. The only health-domain meta-analysis (102
RCTs in children, N = 45,998) finds activity interventions cut
sedentary time by 0.95% of wear time and move screen time and sleep not
at all. In the environmental domain, where spillover is most studied,
the pooled effect on actual behaviour is d =
−0.03.
What does carry across runs through motivation, not willpower. In
the one RCT with a mechanism (Mata et al. 2009, N = 239), an
exercise programme improved eating self-regulation and the effect was
fully mediated by autonomous motivation. In rehabilitation
patients (N = 470) exercise predicted later healthy eating
through habit strength and the belief "sticking to a healthy diet is
easier when I exercise". In students (N = 322) activity
predicted fruit and vegetables at β = 0.20 via self-efficacy. These are
prospective associations with plausible mediators — Moderate — not
cascades.
The two economics trials that use the phrase test something else.
Bjorvatn et al. (2021, ~700 unemployed Norwegian youth) set goals for
sleep, exercise and substance use as a bundle and report better
employment a year on; it does not test whether one habit spread to the
others. Cappelen et al. (2025) gave students a gym card and found fewer
dropped classes — abstract only.
Negative spillover is real, conditional, and smaller than
feared.
Private moral licensing — "I exercised, so I can" — is about zero
once publication bias is removed: 115 experiments, N = 21,770,
give g = 0.65 when the good deed was observed and a
corrected g = −0.01 when it was not (Rotella et al.
2025). The earlier d = 0.31 (Blanken et al. 2015) was mostly
publication bias.
Exercise does not reliably drive overeating: acute-exercise
meta-analysis ES = 0.14, training studies at most +102 kcal a day.
Compensation appears when exercise is framed as "fat-burning" and only
in low-regulation, tired exercisers (N = 96). The much-quoted
"just thinking about exercise makes me serve more food" carries a 2019
corrigendum and a co-author who resigned over research misconduct; do
not cite it.
In daily life both happen. A week of experience sampling (N
= 235, five prompts a day) found healthy → healthy consistency
and balancing, and found compensatory health beliefs unrelated
to actual licensing.
The one cascade with an effect size is local. A
preparatory act — packing the bag, laying out the mat — cues the target
behaviour: preparatory habit predicted exercise at β = .20 (N =
181) where execution habit did not, and teaching cue use and consistency
to new gym members raised accelerometer-measured activity by d =
0.39 at eight weeks (N = 94). Singh et al. list
preparatory habits among the determinants of habit strength (§9). This
is the keystone idea shrunk to its evidenced size: a small act that
reliably starts a larger one, within the same behaviour.
Identity is the plausible route across behaviours, and it
nets to nothing on its own. Habit and identity correlate at
r = 0.55 across 19 studies and 13,340 people — Strong
as an association, mostly cross-sectional. But reminding people of a
past good behaviour raises identity and lowers guilt, and the
two cancel: net spillover ≈ 0 (Lacasse 2016, N = 377 and 172;
replicated 2021). Only an explicit label ("environmentalist") kept the
identity gain without the guilt drop. A field experiment adopting one
new behaviour for three weeks (N = 125) found spillover
"limited and relatively small".
For Cultivar.
Do not build a keystone. No habit gets marked as the one that
unlocks the others, and the app never says "start with this and the rest
will follow". There is no mechanism left for it to be true, and the
message is the coach's voice.
The Together card is the n = 1 version of this question and
it should stay descriptive. Two habits that rise and fall together are
consistent with shared context (§3), shared mood, a preparatory chain,
or coincidence, and the experience-sampling data say consistency days
and balancing days both exist. "These two move together" is honest;
"this one carries that one" is a claim the evidence does not make for
populations and the app cannot make for one person from co-occurrence
alone.
Where a cascade is real it works by starting, and the app
already leans that way. The smaller-ask advice (§8) and the
preparatory-habit finding are the same idea from two directions: the
thing that cues the thing is worth more than the thing.
The mediators that survive — enjoyment, autonomous motivation,
self-efficacy from mastery — are the ones §9 chose to record (affect)
and to leave unasked (self-efficacy). Nothing to change.
Identity spills over only when it is labelled, and labelling is the
app's least favourite voice. The record is the identity; leave the label
to the person.
What this means
the reflection should look for
The reflection has been producing restatements because the summary
gives it aggregates and the prompt asks for "an observation". The
research above names the observations that would actually be worth
making. This is the catalogue to build toward.
#
Pattern
Grounded in
Data needed
Status
1
A day asks for more than it is getting
§8 consistency ceiling
weekday load + per-habit rate on that day
built
2
A habit is kept at a scattered hour vs a consistent one
§3 context stability
completedAt fitted per habit
(domain/hours: von Mises mixture on the last 60 days,
change points over everything); two settled hours read as a routine, a
moved hour reads as a move with its before and after
built
3
A gap is approaching a week, not just a day
§2 Lally vs Armitage
consecutive missed due days
partly — dated gap spans are sent; the wilting signal still fires at
3 days
4
A demanding habit is slower to settle, and that is normal
§8 complexity
habit age + target/unit as a proxy for effort
built (age, target, note all sent)
5
Two habits rise and fall together
§3 shared context
per-day co-occurrence between roots
built — pairs ranked by distance from
independence
6
A habit went quiet when something else changed
§3 habit discontinuity
change-point against schedule edits
partly — dated gaps are sent, schedule edits are not
7
Early repetitions are buying more than later ones
§1 asymptotic curve
habit age
built
8
One habit is kept more on days another came first
§10 instigation, §11 preparatory chain
completedAt order within a pair
not built — Together uses co-occurrence, not order
Verified live against Claude Opus 5 once the data was
there. Before, the year window said "Reading is the only
thing due on Saturday, and that day sits at 50 percent" — a figure
read back. After: "Stretch reads as the habit kept least often, but
on Tuesday and Saturday it runs at 60 percent, the same rate meditation
holds on six days of the week; the 40 percent comes from Thursday
alone." That is an aggregate being corrected by its own breakdown,
which is the thing the five rule-based cards structurally cannot do.
The distinction the app keeps failing to draw: an observation
restates a number in prose; a derived insight proposes a
mechanism. "Thursday sits at 50%" is the former. "Thursday is
the only day carrying four habits, and the two that fail there are the
two that need the most setup" is the latter. Only the second gives the
person anything to act on, and only the second earns the cost of the
call.
What the research does
not support
Worth writing down, because these are the things that will keep being
suggested:
"21 days to form a habit." No evidential basis. The
median is 66 and the range reaches 254.
Habit formation as a countable, completable
process. The curve is asymptotic; there is no day on which it
is done.
Treating a single miss as meaningful. It costs 0.29
points and nothing lasting.
Anchoring to a routine as a faster route than any other
consistent cue. The one RCT (Keller 2021) found routine and
clock-time cues interchangeable, and the one positive study did not
replicate. What is supported is a consistent cue enacted daily — which
is §3.
Keystone habits. The mechanism Duhigg gave them
(trainable willpower) did not replicate, the spillovers that do exist
are small and fade within a year, and habit science has never adopted
the term. No habit unlocks the others.
Streak counts as motivation. The evidence supports
removing the cliff, not displaying the number.
Celebration as reinforcement. Lally provided no
extrinsic reward and habits formed anyway; the participants chose
behaviours they found intrinsically rewarding. Fogg's "celebrate
immediately" is a framework claim, not a finding.
Open questions
Things this app is in an unusually good position to answer for its
own user, and currently cannot:
Does consistency of hour predict this person's rate, as §3
predicts? Needs completedAt read, which is one query
away.
Is three missed due days the right threshold for the wilting signal,
or should it track the habit's own rhythm more closely? §2 suggests a
week is where the evidence bites.
Does the app's own score correlate with anything a person would
recognise as the habit getting easier? Cultivar measures compliance; the
literature measures automaticity. They are not the same thing, and the
gap is worth remembering whenever a number here is treated as
progress.
Does the order of two habits within a day matter for this person?
Judah 2013 said after beats before and Judah 2018 said it makes no
difference; the app has completedAt for both halves of
every pair and could tell.