66 days
Median time for a new behavior to become automatic (Lally et al., UCL)
20 hours
Deliberate practice needed for functional competence in a new skill
90 days
Default outer bound for a career experiment to produce usable signal
51%
US employees who are disengaged at work (Gallup)
By the end of this guide, you will have a working career experiment calendar: two or three bounded tests, each with a hypothesis, a start date, an observation window, a decision gate, and a kill date. No vague "I should probably look into that" energy. No five-year plans written in disappearing ink.
Here is the uncomfortable truth most career advice skips over: you cannot think your way into a better career. You can only test your way there. Psychologist Herminia Ibarra at London Business School has documented for two decades that people change careers by running small experiments in the real world -- not by introspecting until a lightning bolt hits. In "The Authenticity Paradox," Ibarra describes career transition as a repeating cycle of test, learn, adapt, test again. The one ingredient missing from almost everyone's version of that cycle is a clock.
Without a timeline, a career experiment quietly becomes a hobby. With a timeline, it becomes evidence you can make a decision on. The difference between the two is roughly 90 days and one written kill criterion.
What makes this guide different
Most career experiment advice stops at the idea. This guide is about the schedule -- the durations, gates, and pre-committed exit conditions that convert curiosity into a decision. Grab a calendar app before you keep reading. You will need it by Step 5.
Prerequisites: What You Need Before You Set a Single Date
Do not skip this section. Every failed career experiment I have watched -- including three of my own -- failed at the prerequisites stage, not the execution stage. You need five things before Step 1.
- A question, not a conclusion. "I want to become a UX designer" is a conclusion. "Do I enjoy the repetitive parts of UX work more than the repetitive parts of my current job?" is a question. Experiments can only run against questions.
- Four to six protected hours per week. Not twenty. Experiments that demand twenty hours collapse the first week your day job gets busy. Research on skill acquisition summarized at JoshKaufman.org suggests the first 20 hours of deliberate practice are where the steepest gains live -- but only when those hours are consistent rather than heroic.
- A written baseline. Before day one, record where you are right now: satisfaction (1-10), income, hours worked, and your energy level on a normal Tuesday afternoon. Without a baseline, you cannot detect change later. You will just argue with yourself.
- A decision rule you wrote in advance. What result makes you continue? What result makes you stop? Write both down before you start. This single habit is the difference between a career experiment and a very organized coping mechanism.
- A way to score your starting position. Before you commit 90 days to a direction, it helps to know how exposed your current career actually is. Try the free Career Pulse Score at Workings.me -- it takes a few minutes, gives you a future-proofing baseline, and you can re-run it at the end of your experiment to see whether the test actually moved anything.
Step 1: Write Your Hypothesis in Falsifiable Form
Why this step matters: "Explore product management" cannot be evaluated. An unevaluable experiment runs forever, which means it never produces a decision, which means it was never really an experiment.
How to execute: Use this template and fill it in literally.
The hypothesis template
"If I spend [X hours per week] doing [specific activity] for [Y weeks], then I will be able to answer [specific question] with a yes, a no, or a not-yet."
Weak: "Learn data science."
Strong: "If I spend five hours a week for eight weeks cleaning and visualizing public datasets in Python, then I will know whether the messy-data debugging part of this work energizes or drains me."
Common mistake: Building the hypothesis around achievement instead of fit. Finishing a Coursera specialization tells you that you finish courses. It tells you nothing about whether you want the job. Make the question about the daily texture of the work -- the Tuesday afternoon of it, not the graduation photo.
Step 2: Match Your Experiment Type to Your Question
Why this step matters: Different questions need different observation windows. Running a two-week skill experiment is useless because nothing has had time to get boring. Running a six-month series of coffee chats is not an experiment -- it is procrastination with good manners.
How to execute: Pick the row that matches your question, then use the default window as your starting point.
| Experiment type | Answers the question | Default window |
|---|---|---|
| Skill test | "Do I like doing this work?" | 8-12 weeks |
| Shadow project | "Can I deliver in this context?" | 4-6 weeks |
| Exposure and network | "Is this world what I think it is?" | 2-4 weeks |
| Earned-revenue test | "Will someone pay for this?" | 6-10 weeks |
Common mistake: Using the cheapest experiment type for the most expensive decision. Three informational interviews is not a test of whether you should quit your job. It is a vibe check. If the decision costs you a year of income, the experiment needs to cost you at least six weeks of evenings.
Step 3: Apply the 30/60/90 Rule to Set Your Observation Window
Why this step matters: Week one of anything feels amazing or terrible for reasons that have nothing to do with the work. You need enough time for the novelty effect to wear off and the real texture to show up.
How to execute: Split any experiment longer than six weeks into three phases.
- Days 0-30 -- Novelty phase. Everything is new, your brain is flooded with dopamine, and your judgment is unreliable. Do not make a decision in this phase. Keep showing up.
- Days 31-60 -- Reality phase. The interesting part. Boredom, friction, and genuine satisfaction both surface here. This is where your daily log starts earning its keep.
- Days 61-90 -- Signal phase. You now know what this work actually feels like on an ordinary week. You can trust your read.
This maps neatly onto the habit research. Phillippa Lally and colleagues at University College London found a median of 66 days for a new behavior to become automatic, with a range from 18 to 254 days. Sixty-six days sits almost exactly at the two-thirds mark of a 90-day window -- the point where the behavior stops requiring willpower and starts revealing preference.
Common mistake: Quitting at day 12 and calling it data. Day 12 measures your tolerance for being bad at something, not your fit for the work.
Pro tip: expect the week-three dip
Nearly every experiment has a wall around week three when the initial novelty fades but competence has not arrived yet. Pre-commit now: "I will not evaluate this experiment before day 30." Write it in your calendar as an event. That one pre-commitment saves more careers than any personality test.
Step 4: Pre-Register Your Evidence Standard
Why this step matters: Humans are spectacularly good at deciding what counts as success after seeing the results. Borrowing the scientific practice of pre-registration -- deciding your criteria before collecting data -- closes that loophole. Research on the planning fallacy, summarized in this HBR piece by Lovallo and Kahneman, shows we consistently misjudge our own future effort and outcomes. Pre-registration is the cheapest known antidote.
How to execute: Define three metrics before day one and log them daily in a simple note or spreadsheet.
- Effort metric: Did I complete my planned hours today? Yes or no.
- Enjoyment metric: How did the work itself feel? 1-5.
- Evidence metric: One concrete thing I produced or learned. A file, a call, a paragraph, a shipped thing.
Then write your thresholds. Example: "Continue if my average enjoyment is 3.5 or higher over weeks 5-12 and I produced at least eight pieces of evidence. Stop if enjoyment averages below 2.5 for two consecutive weeks."
Common mistake: Using only output metrics. If you log "finished the module" you will learn nothing, because you can finish anything out of stubbornness. The enjoyment metric is what actually predicts whether you will still be doing this in three years.
The Tuesday test
Here is a faster version of the enjoyment metric. At the end of each session ask: "If this exact task were my job, and it was a rainy Tuesday in November, would I still do it?" Your gut answer is data. Write it down.
Step 5: Build the Time Budget and Block the Calendar
Why this step matters: Career experiments do not die from lack of interest. They die from lack of a slot. Energy is the real constraint, and energy is highest in the morning for most people, which is exactly when we give our best hours to an employer.
How to execute: Structure your four to six weekly hours into three distinct block types.
- Deep block (2 hours, once or twice weekly). The actual skill work. Protect this like a meeting with your boss. Put it on Google Calendar as a recurring event marked busy. If you want it scheduled automatically around your existing commitments, tools like Reclaim.ai or Motion will defend the time for you.
- Shallow block (1 hour weekly). Reading, browsing job postings in the target field, watching one tutorial. Low stakes, keeps momentum.
- Social block (1 hour weekly). One conversation with someone already doing the work. This is the block everyone skips and the one that most reliably changes outcomes.
Log the experiment in a single note -- Notion and Obsidian both work fine. One page, one header with your start date and kill date, and a dated log beneath it. That is the entire infrastructure requirement.
Common mistake: Booking six hours on Sunday. Sunday experiments die by week three, because Sunday is the day life happens. Two weekday evenings and one Saturday morning beats one heroic weekend block every time.
"I had been 'thinking about' moving from agency account management into product for two years. Three notebooks, zero decisions. The thing that finally worked was the deadline. I gave myself a 10-week experiment -- two evenings a week writing specs for a fake product, plus one Friday call with a PM at a company I admired. Around week 6 I noticed I was looking forward to the spec work more than my actual job, which was not the answer I expected. My kill criterion was 'if I dread it by week 8, I stop.' I did not stop. I moved into a product role 14 months later and I still use the same 90-day format for every big career question now."
Step 6: Add Decision Gates and Kill Criteria
Why this step matters: The single most valuable line in your experiment plan is the one that tells you when to stop. Without it, you will either quit too early out of frustration or drag a dead experiment for a year out of sunk-cost loyalty. Neither produces a decision.
How to execute: Schedule three calendar events with hard agendas.
- Day 30 gate -- Do not decide. Review your log for completeness only. If you completed fewer than 70 percent of your planned sessions, the problem is your schedule, not the career question. Fix the calendar and continue.
- Day 60 gate -- Provisional read. Compare your average enjoyment against your pre-registered thresholds. If you are below the stop threshold, stop. Do not negotiate with the data you wrote yourself.
- Day 90 gate -- Decide and document. Write a one-page verdict: continue, scale, pivot, or stop. Then write the next experiment's hypothesis before you close the file.
Common mistake: Building a single end-of-experiment review instead of gates. A single gate at day 90 means you cannot catch a scheduling problem until the experiment is already dead. Gates exist to catch problems in time to fix them.
Step 7: Instrument the Experiment Without Turning It Into a Job
Why this step matters: Measurement should take under three minutes a day. If your tracking system takes longer than your actual practice, you have built a second job and you will quit the first one you can defend quitting.
How to execute: Three tools, nothing more.
- Daily 60-second log. A single note with the date, hours completed, enjoyment score 1-5, and one line of evidence. That is it.
- Weekly 10-minute review. Every Sunday, read the week's entries and write one sentence: "What did this week tell me?" Twelve sentences over 90 days is an astonishing amount of self-knowledge.
- Monthly external signal. One piece of outside feedback per month -- a portfolio critique, a paid micro-gig, a colleague's honest read on your output. Internal enjoyment tells you what you like. External signal tells you whether the market agrees.
If you want to track whether the experiment is actually improving your long-term position, re-run your Career Pulse Score at the 90-day mark and compare it to your baseline. Future-proofing, like fitness, is only measurable against a starting point.
Common mistake: Over-instrumenting. Ten metrics means zero metrics, because you will stop filling them in by week two. Three numbers, logged every day, beats a beautiful dashboard you abandon.
Step 8: Run the Post-Mortem and Stack the Next Experiment
Why this step matters: A single experiment rarely changes a career. Three stacked experiments over nine months almost always does, because each one narrows the search space. The post-mortem is where that narrowing happens.
How to execute: Answer five questions in writing, in under 30 minutes.
- What did I actually do? (hours completed, evidence produced)
- What did I learn about the work, not just about myself?
- What did I learn about the context -- the industry, the pay, the people?
- What surprised me?
- What is the next smallest test that resolves the biggest remaining uncertainty?
Then pick your next experiment from one of three moves. Scale if the signal was positive and the bottleneck is now depth. Pivot if you liked the domain but hated the task type -- this is the most common outcome and the most useful one. Stop if the answer was a clean no, and treat that as a win, because a clean no saves you three years. You can see how fast industries are reshuffling in the McKinsey Future of Work research -- the more volatile the target field, the shorter your experiments should be.
Three Timeline Scenarios You Can Copy
The Weekend Skeptic (4 weeks, 6 hours per week). Best for a low-stakes curiosity. One 2-hour deep block Saturday, one 1-hour research block Wednesday, one 30-minute call each week. Gate at day 28. Use this to rule things out fast.
The Standard 90 (12 weeks, 5 hours per week). Best for a genuine career direction question. Two 2-hour deep blocks, one 1-hour social block. Gates at day 30, 60, and 90. This is the workhorse format and the one most readers should default to.
The Earned-Revenue Sprint (8 weeks, 8 hours per week). Best for testing freelance or consulting viability. Weeks 1-2 define an offer, weeks 3-6 pitch to ten real prospects, weeks 7-8 deliver one paid engagement. Gate at day 56 with a binary question: did anyone pay me? Nothing else counts as signal. The BLS Occupational Outlook Handbook is a useful reality check on typical earnings and growth before you commit hours to a field.
Common Timeline Mistakes That Kill Experiments
- No kill date. An experiment without an end date is a mood. Set the date on day one.
- Deciding in the novelty phase. Weeks one through four are chemically unreliable. Wait for the dip.
- Treating busyness as progress. Forty hours of tutorials with zero evidence artifacts is not an experiment, it is consumption.
- Running four experiments at once. You will do all four badly. Two is the practical ceiling; one is better.
- Ignoring the context tests. Enjoying the work is only half the answer. You also need to know whether the pay, hours, and people are acceptable. Add one context metric to your log.
- Repeating the same experiment. If experiment two looks identical to experiment one, you have not extracted the learning. Each new experiment should resolve a different uncertainty.
Insider tip: the evidence artifact rule
Every week of your experiment should produce one artifact that exists outside your head -- a file, a post, a spec, a spreadsheet, a recorded call, a small paid deliverable. Artifacts are the only thing that survives the experiment. Six months later you will not remember how you felt in week four, but you will still have the artifact, and it will be the thing that gets you the interview.
Quick-Start Checklist
Run this before day one
- [ ] One falsifiable hypothesis written in the "If I... then I will know..." format
- [ ] One experiment type selected from the four-row table, with its default window
- [ ] Baseline numbers recorded: satisfaction, income, hours, Tuesday energy
- [ ] Career Pulse Score run once for a future-proofing baseline
- [ ] Three metrics defined: effort, enjoyment, evidence
- [ ] Continue threshold and stop threshold written down, not remembered
- [ ] Four to six weekly hours blocked in the calendar as recurring events
- [ ] Day 30, 60, and 90 gates scheduled with agendas
- [ ] Kill date entered in the calendar with a reminder one week before
- [ ] Post-mortem template created and saved for day 90
That is the whole system. It is not complicated, but it is unforgiving of vagueness. The people who get somewhere interesting are almost never the ones with the best career plan. They are the ones who ran the most experiments and had the discipline to end each one on schedule. Set the dates, protect the hours, honor the kill criteria, and let 90 days do the work that two years of thinking never will.