Case study
TidyDay
TidyDay — a day planner for overloaded students that helps you decide what to do when your day stops fitting.
My role: everything. Research, product strategy, interaction design, conversation design.
2The short story
Every scheduling app can build you a plan. I got interested in what happens after the plan breaks — because that's most days, and it's where every product I tested falls apart. The market leader's answer to an impossible day was scheduling a two-hour problem set at 4:45am. And reporting no problem.
TidyDay is my answer: a planner that repairs your broken day by comparing what each fix would actually cost you, tells you honestly when nothing free is left, and shows you everything it has learned about you in a list you can edit. This case study covers how I got there — including one full concept I had to kill, four of my own claims that died under evidence, and the research I chose not to fake.
This project started as something else. In 2025 I built a concept for an AI assistant that organizes your day for you — and the evidence told me, twice and eighteen months apart, that nobody needed it. That evidence-led pivot is where this case study really begins, so I've kept it in instead of hiding it.
What survived the pivot is a sharper question: not "how do I generate a schedule" — that's arithmetic, and seven shipped products already do it — but what should software do at 10:20am when the lecture ran long and the day no longer fits? Existing tools reshuffle silently, cram work into the middle of the night, or report failure and walk away. None of them will say the one honest thing: something has to go.
The result is a mobile app design built on five ideas: repairs are compared by cost, not applied first-come; refusing to produce a fake plan is a valid outcome; the app opens the conversation instead of waiting to be opened; every sacrifice it proposes gets followed up on; and everything it believes about you sits on one screen you can read, correct, or delete.
3The pivot — killing my own concept
Users rejected my first version of this product eighteen months before my own market research did. Four survey respondents said it had no differentiator; one already owned it. When my 2026 competitor teardown reached the same conclusion from the opposite direction, I stopped treating the pivot as a choice. The evidence forced it — and that's a better story than pretending I was right the first time.
The 2025 concept was "a personal assistant that organizes your day for you." I surveyed it — 18 responses, Jan–Feb 2025 — and built interview-based personas for it as coursework. The survey came back blunt: four people independently said it had no differentiator, one said "I have it" about an existing tool, and a third of respondents said the tool wasn't the problem — they ignore reminders anyway.
In 2026 I tore down seven competitors (Motion, Reclaim, Morgen, Structured, Sunsama, Akiflow, Apple/Google Calendar) and reached the same verdict independently: generating and reshuffling a schedule is a crowded commodity. Motion and Reclaim auto-schedule. Apple and Google have shipped "time to leave" since 2012. A new entrant doing the same thing with a new colour palette answers no question a hiring manager or a user would ask.
Two independent lines of evidence, opposite directions, same conclusion. So I pivoted the thesis from generating the plan to negotiating the break — and kept every artifact of the dead concept in the project as evidence, including my own 2025 personas, none of which survived the new framing. Documented, dated, checkable.
One thing the old survey did hand me for free: two respondents, unprompted, asked for honesty from a product that knows things about you. "Make it honest" was one full response. That instinct became the memory screen — the feature the whole design now leans on.
4Research — testing the market for real
I built a deliberately impossible week — 17.5 hours of tasks against 10 hours of free time — and drove the two market leaders through it with a real account and a real credit card. Then I asked Motion's AI chat to make the exact trade my product proposes. It replied "Perfect! ✅" — and documented its own failure three lines later. Four of my market claims died during this research. The claim that survived is stronger than the one I started with.
Desk research told me what products say they do. I wanted to see what they actually do when a week doesn't fit, so I staged one: a commuter student's week, 58 calendar events generated by a script I wrote, ~70 minutes of commute each way, a weekend shift, a daily gym habit, a weekly club — and ten tasks totalling about 17.5 hours against roughly ten hours of real slack. Deliberately unsatisfiable. Then I ran Motion and Reclaim against it, hands-on.
What I observed, first-hand:
- Motion scheduled a two-hour problem set at 4:45–6:45am and reported no problem. Not an isolated slot — group project work at 22:15–00:15, seminar reading at 23:15–01:15. When a day is full, the small hours are where Motion puts your work.
- It scheduled two tasks past their own due dates — eleven and fourteen hours late — with no flag, while its documentation promises "the next best slot before the deadline." It spent the night and missed the deadline anyway.
- It stated failures without remedies. Three days running it said lunch and recurring tasks couldn't fit — then left them stacked on the calendar and moved on. The conflict badge sits among enough icons that I missed it while actively looking for it.
- Other tasks were silently pushed to next week. No notification. I found out by looking.
- The best finding came from its chat. I typed the exact sacrifice my product is built around: "remove gym/workouts and reschedule." Motion replied "Perfect! ✅ All critical academic tasks have been rescheduled" — above a table marking four of six tasks past due — then admitted "Delete your calendar events manually — I cannot do this." Gym is an event. Work is a task. The trade crosses Motion's two data models, and the product has no term for it. That's not a UI bug; it's structural, and it became a design decision in TidyDay.
- Reclaim's "Issues" panel does propose a fix — postponement, with the same template time slots for every conflicting task, onto times I was already busy. A proposal that ignores what it costs and where it lands isn't help; it's confident wrong advice.
- Neither product has any screen showing what it learned about you. I searched Motion deliberately — settings, profile, task defaults, analytics. Nothing. The one gap that survived every research pass, confirmed first-hand.
Four of my early claims did not survive this: "competitors don't notice," "nobody warns you," "no competitor has chat," "nobody asks" — all checkably false, all banned from the project. The corrected claim is narrower and harder to attack: every piece of the interaction already ships somewhere — detection, failure notices, chat, approval steps, end-of-day rituals. No product connects any of it to the moment the day breaks, and none of them will tell you something has to go.
I also ran a screening survey into the target population. Three responses so far, all transit commuters at 50–75+ minutes each way. Days break one to three times a month, not weekly. And the sharpest single data point in the project: asked what they'd never skip, they answered talking to family, leisure time, praying — three things that appear on no calendar and that no one would ever enter into a scheduling app.
What I didn't do — on purpose. I planned five-plus user interviews and could not recruit them: the screener converted three respondents into zero conversations, and campus community approvals were pending indefinitely. I had a 2025 dataset and old personas I could have quietly counted as user research. I didn't. The study is declared evidence-led about the market, assumption-led about depth of user understanding — every requirement carries its evidence source, and every assumption sits in a register that doubles as the interview backlog. A reviewer should be able to check any claim I make in one search. Several of mine failed that test during this project; the ones that remain have survived it.
5The problem
An overloaded student's day breaks one to three times a month. When it does, their tools produce technically valid, humanly useless repairs — 4:45am work blocks, template postponements onto busy slots, silence. So the student absorbs the overflow into the night, or drops something and carries the cost alone. One respondent dropped a side project on a broken day, never made it up, and weeks later still answered "was that the right call?" with: "Not sure."
The user: a post-secondary student with immovable anchors — fixed class times or hard deadlines — more commitments than hours, and a droppable set whose cost lands on them personally. Often a transit commuter, which matters twice: the commute supplies genuinely fixed points to plan around, and it's where the market's "time to leave" feature is weakest — Apple's is driving-biased, and this population rides the bus.
Who it's deliberately not for: knowledge workers (their day is meetings, and meetings move), and parents (their droppable set is relational — the cost of skipping school pickup lands on someone else, and "skip the recital, Thursday is free" is not a negotiation an app gets to have). Naming who you're not building for is what keeps assumptions from quietly widening to cover everyone.
This population is basically un-tooled. Of my three screener respondents, two use a phone calendar, one uses memory, and none has ever had an app move something without asking. They aren't switching from a competitor. They're adopting a first planner — which makes onboarding weight and first-week trust the highest-risk surfaces in the design.
The problem statement I'm designing against:
When an oversubscribed student's day stops fitting, existing planners produce technically valid, humanly useless repairs — uncosted, destination-blind, or silent — and none will say that something has to go. The student absorbs the overflow into the night or drops something and carries the unresolved cost alone.
And the hardest sub-problem, straight from the data: the things people protect most fiercely aren't in the calendar at all. An app that proposes sacrifices while blind to the three things this sample would never sacrifice isn't just wrong — it's insulting. That problem shaped one of the design decisions I'm most confident in (section 6).
6The design answer
Five decisions carry the product, and each one exists because I watched the alternative fail: repairs are costed and compared instead of applied first-come; "this day doesn't fit" is a legitimate answer; the app speaks first, and never states a problem without a proposal; every sacrifice it proposes gets followed up on; and everything it believes about you is one readable, editable screen. The design work isn't the schedule — it's what the software does when the schedule dies.
One kind of thing. Motion couldn't accept my sacrifice because gym was an "event" and work was a "task" — two tables, no trade between them. Reclaim has the same wall from the other side. So TidyDay holds one commitment model: a lecture, a gym session, an assignment and a shift are the same type with different attributes. Any commitment can be moved, split, proposed for dropping, or refused — and a trade can happen between any two of them. The market's structural gap became my data model.
Repair is a costed comparison, not a fallback chain. My first version of this was a ladder — reschedule, then split, then escalate. I killed my own design here: a ladder takes the first fix that works, and the first fix can be the worst one. Moving a four-hour task to tomorrow can wreck tomorrow worse than losing 45 minutes tonight — but a ladder never compares them. So TidyDay generates candidate repairs — reschedule, swap, split, shorten, ask, renegotiate, spend protected time, drop, refuse — costs each against what it knows about you, applies the free ones (and says so), and proposes the costly ones as two or three options with the tradeoff visible. Never a menu of everything. A recommendation, first.
Refusing is an outcome. Sleep is not capacity. The 4:45am block exists because Motion's arithmetic always has an answer. TidyDay's doesn't: sleep and any hours you mark protected are simply not schedulable space, by default, no setup required — because the user drowning in a broken week is not the user who carefully configured quiet hours. When nothing acceptable remains, the product says the day doesn't fit and asks. One refinement I care about: the app may offer protected time as an explicit, costed, refusable option — "this takes 45 minutes from tonight's sleep" — because my own data shows students already make that trade unaided, at 2am, untracked. It may never spend it silently. You pay protected hours knowingly or not at all.
The app speaks first, and never empty-handed. Every conversational surface in this market waits to be opened. TidyDay's chat starts because something happened in your day — and every message it sends carries a proposal, an action it's already taken, or an honest "I'm out of options, what gives?" Never a bare warning. I watched Motion name a problem and walk away; a notice with no remedy is indistinguishable from noise, and gets treated as noise. Suggested replies are real messages ("already half done", "gym matters today") so answering takes one tap at 10:20am with one hand — and free text is always there, because "I've done half already, I need 30 more minutes" corrects the app's premise and hands it a duration estimate in the same sentence.
A sacrifice must close its loop. The most trust-destroying thing this product could do isn't a wrong estimate — it's asking you to give something up and then losing the thing you gave it up for. So every accepted sacrifice gets one follow-up at the calm end of the day: did it pay off? If it didn't, the app owns that plainly, shows what it learned, and never proposes the same trade again without referencing the loss. My screener data has a person stuck in exactly this loop with no help — dropped it, never made it up, still not sure it was right.
The memory screen. Everything TidyDay believes about you — "you said missing gym is annoying but fine, and you've chosen it over club twice" — lives on one screen, in plain language, with the evidence behind each belief linked to the conversation that produced it. Editable, deletable. No competitor has this surface; it's the one gap that survived every research pass. It's also what makes an app that judges your priorities tolerable at all: a wrong suggestion you can see the reasoning for is correctable; a wrong suggestion from behind glass gets the app deleted.
Protection without disclosure. The untouchable things — family, leisure, prayer — aren't calendar entries, and asking users to declare them would force the most sensitive categories of a life into a scheduling database. So a protected block needs no label. Mark the time; never say why. The scheduler knows that 6–7am is untouchable, never what it is. The app can't schedule over what it protects, and can't leak what it never knew.
And the everyday planner is the sensor. Most days, nothing breaks — TidyDay is just a planner with travel-aware buffers. But every ordinary interaction is quietly load-bearing: adding a commitment asks one three-option question ("if you had to miss this — fine / annoying / not an option?") that seeds the importance model; recurrence gives deferral cost for free; completion history teaches real durations, because Motion's fixed duration dropdown forced me to enter a three-hour task as four and then planned confidently against a number I knew was wrong. The ordinary days are where the model — and the trust — get built, visibly. Motion auto-created tasks from beliefs I never saw. TidyDay's collection happens in front of you.
What I deliberately cut: a web companion (paid for the memory screen instead), every suggestion category beyond three designed instances, and any AI labelling on things that are actually arithmetic — "leave by 8:10" is subtraction and an API call, and calling it AI would be marketing dressed as capability.
7Design & testing — in progressIn progress
Running now: lo-fi screens for the five surfaces that carry the product — the day view, the negotiation chat, the memory screen, adding a commitment, and the moment the app asks permission to speak. Four-plus test sessions with before/after preserved, then hi-fi and a clickable prototype of the full loop: day breaks → conversation → pushback → plan changes → memory screen shows what it concluded.
(This section fills in as the design phase lands. Structure ready: lo-fi rationale per screen · what the test sessions changed, with befores kept · conversation design — the app's voice as a written spec, not an afterthought · hi-fi and the prototype walkthrough.)
8What I'd measure
The key metric is one I refuse to maximize. Proposal acceptance should sit in a healthy band — around 40–70% — because near-100% means the user is rubber-stamping and the "asking" is theatre, and near-zero means the model is mistrusted. An app optimizing acceptance would learn to propose only what you already wanted to hear. That refusal is the product's whole point, in metric form.
Product metrics, if this shipped:
- Proposal acceptance rate, as a band (≈40–70%), not a target to push up. The reasoning above — this is the one number that keeps the product honest.
- Refusal frequency — rare, watched for trend. "This day doesn't fit" should be the exception. If it trends up, my capacity model is wrong about how often days truly break, and that gets reported, not tuned away.
- Memory-screen corrections — some, but not mass deletion. Zero corrections means nobody reads the screen. Mass deletion means it alarms instead of reassures. Either tail is a finding.
- Retention through the first proposed sacrifice. The first "you might have to skip gym" is the trust cliff. Users retained past it, versus retention generally, is the single number that says whether the defining interaction works.
- Guardrails at zero, by construction: work placed in protected hours without an accepted offer; a lost sacrifice re-proposed without acknowledging the loss; a failure message with no proposal attached; a success summary its own details contradict. I watched shipped products do all four.
What I deliberately would not measure: time-in-app (a planner that fixes your day faster wins while that number falls), tasks completed per day (this product proposes doing less, on purpose), and streaks — a streak is a nag with confetti, and the product's register is calm honesty.
For the case study itself, the measures are humbler and real: task completion across the lo-fi sessions, and one comprehension check I care about most — after seeing the memory screen, can a participant tell me, unprompted, what the app believes about them and why.
9Limitations — what this study doesn't know
I'll say it before a reviewer does: no user interviews happened. I recruited, the funnel produced three respondents and zero conversations, and I chose declared assumptions over faked findings. The register of what I'm assuming — and what would test each assumption — is part of the work. The biggest open question could still invalidate the thesis: nobody has yet told me whether they'd accept advice on what to give up from software.
- Depth of user understanding is assumed, not evidenced. The market behaviour is observed first-hand; the users are reached only by a three-response screener and published commuter research. Every requirement carries its evidence tag, n=3 is written wherever the screener is cited, and the assumptions register doubles as the interview backlog — interviews landing later upgrade entries in place.
- The core premise is untested. Would people accept software proposing what to drop — and if not, is the objection to being advised at all, or to the advisor being a machine? My discussion guide has two sections built to break this. They haven't run. If the answer comes back "that's mine" or "not from an app," the thesis reopens, and I'd rather learn that at lo-fi than after hi-fi.
- Asking may not beat being managed — for everyone. The ADHD-planning literature records the same population crediting auto-scheduling with removing real decision paralysis and describing the loss of control as suffocating. Both are true. That's a segmentation boundary, and finding it is the droppability model seen from the other side.
- The escalation is rare. One to three broken days a month in my sample, and mostly a shuffle fixes them. I've designed for that honestly — the everyday planner carries the daily value — but the case study's centrepiece is, by the data, a monthly event. The number stays in the writeup.
- Some things stay invisible. Protection without disclosure still can't protect what was never entered. The app can hold an unlabeled block; it can't know the block should exist.
- One research finding rests on recall (Reclaim's template postponement slots — no screenshot), and my trial calendars weren't isolated from each other, so I make no claims comparing the two products' scheduling quality. Product-internal findings stand.
- Privacy is a real surface, named. A model of what you'll give up under pressure is sensitive by nature. The memory screen and unlabeled protection are the mitigations; the limitation doesn't vanish because I designed around part of it.
10What this project taught me
Absence claims are the cheapest sentences to write and the most expensive to be wrong about. I wrote four — "nobody notices," "nobody warns you," "no competitor has chat," "nobody asks" — and evidence killed all four. Each death made the surviving claim harder to attack. The other lesson cost me three rewritten documents: open the evidence before writing the finding.
- "Nobody has X" fails in one search. Four times I claimed a market absence; four times a product proved me wrong — Morgen asks permission, Motion notices and says so, both trialled products ship a chat, Reclaim proposes remedies. The claim that survived is narrower and observed: nobody connects any of it to the moment the day breaks, and nobody costs a remedy. I now treat every absence claim in my work as unverified until hunted.
- Open the screenshots before writing the finding. My verbal trial notes said "a three-hour task at 4:15am." The screenshot says a two-hour problem set at 4:45. Three documents were written around the wrong number before I opened the image — which also held three findings the notes missed entirely. Evidence first, prose second, always.
- Registers never merge. This project carries four separate participant registers — 2025 survey, 2025 coursework interviews, 2026 screener, future interviews — and no datum crosses between them. It's bookkeeping, and it's also the difference between a study a reviewer can trust and one follow-up question collapsing the whole thing.
- The differentiated claim usually sounds worse than the generic one. "Reschedules your day intelligently" sounds great and six competitors already say it. "Tells you when something has to go" sounds uncomfortable — and it's the empty column in the market. My instinct to soften the uncomfortable framing was worth distrusting, and pushing back on my own framing twice made the product better both times.
- Constraints, declared, are stronger than padding. The honest version of this study — observed market evidence, three screener responses, a signed assumptions register — reads as a designer who knows what they don't know. The padded version reads fine right up until anyone checks a date.