Which hypothesis should the MVP test?

Suppose we are building a service for small repair companies. Jobs arrive in messaging apps, dispatchers assign technicians in a spreadsheet, and customers call for updates. The backlog already contains a portal, a mobile app and a dashboard. We could start building tomorrow. But which problem should the first release solve?

If jobs disappear between conversations, test whether dispatchers will manage real work in one place. If jobs are tracked well but customers keep calling, test whether customers can find a useful status themselves. Mixing both experiments makes it harder to explain why people use the first release.

For this pilot, we choose one question: will dispatchers stop using the parallel spreadsheet? Five companies use the service for four weeks, and we track the share of real jobs they complete in it. Those numbers belong to this example. Whether the sample is useful depends on the decision we want to make and how much the participants differ.

Before building, we watch a dispatcher handle actual work. We ask to see the last lost job, rather than ask whether a future screen sounds useful. Perhaps jobs disappear during shift handover. Then a shared list only helps if responsibility transfers with it. Another inbox alone will not fix that break in the process.

We also separate the user from the buyer. Dispatchers want less switching between chats; managers want fewer lost sales and better visibility of technicians. Employees may love a tool that their manager will not pay for. Both belong in the pilot, although we ask them different questions and expect different evidence.

We choose the riskiest assumption. If the data is available and the interface is straightforward, the main uncertainty may be whether employees will leave the spreadsheet. Another week designing infrastructure will not resolve it. A working path for a small group, followed by observing unprompted use, gets closer to the answer.

We write down possible outcomes before launch. Companies moving real work and returning on their own justify another check. Use that happens only after our reminders needs investigation. Leads who agreed to look but never connected do not count as an active pilot. Adding features is not automatically the right response to each situation.

A baseline matters too. We record lost jobs, shift-handover time and how often customers repeat information before introducing the product. Otherwise every improvement can be attributed to the new tool. Measurement should support the decision we intend to make; it need not become a separate research project with more complexity than the pilot itself.

A small release still needs a complete outcome

The dispatcher must create a job, assign a technician, update its status and close it. If only creation works and the rest stays in a spreadsheet, you are testing a form rather than a shared workflow. Discuss the scope through user actions instead of counting screens.

We can simplify every step: one repair category, one role and email notifications. A technician might report the outcome to a dispatcher instead of using a separate app. The job still gets finished, even though some work happens manually. Cutting scope helps until it removes the action that makes the workflow useful.

We walk through one working day. A customer messages in the morning, the dispatcher creates a job, a technician accepts it and reports completion in the evening. In between, the customer changes the appointment. If the service cannot handle that change, work goes back to chat. Rescheduling belongs in the core path even if the first mockup omitted it.

There is more than one way to support each action. Assignment does not need an elaborate calendar: a technician list and a time field may be enough. Smart matching is unnecessary if the pilot asks whether jobs can move between people reliably. The simpler version preserves the outcome while reducing the implementation.

We identify uncommon events that stop the workflow. An incorrect assignment cannot be dismissed as an edge case if it blocks completion. Bulk history export may wait because it does not affect today's job. Priority follows the effect on the pilot, not just how frequently a feature name appears in interviews.

Not every participant needs a separate application. The dispatcher can use the service, the technician receive a notification and the customer open a status link. Frequent independent work will eventually tell us which role needs its own interface. Three roles do not automatically justify three applications in the first release.

We validate scope by doing the workflow with realistic data. We enter an incorrect address, change the time and finish the job. If an undocumented developer intervention is needed, the path has a gap. Finding it in a prototype costs less than discovering it after inviting the first business to depend on the service.

Manual work must not manufacture the result

The team can import old jobs or send invoices manually. That saves development if the experiment concerns job management rather than independent signup or payment. Explain the assistance to participants and record the time it takes.

If a founder enters every job for the dispatcher, completed jobs do not demonstrate the dispatcher's willingness to use the product. Separate the work done by customers from work done for them. Otherwise the pilot measures your effort rather than adoption.

Ten minutes of daily help per company might be a reasonable learning cost. Two hours for every job imply a different business. Compare the manual effort with the future price and workload before treating completion as commercial validation.

Manual operations need a design too. Who connects a company, receives its data, stores the outcome and answers questions? "A developer will figure it out" creates operational debt before adoption. A short instruction and an owner are enough for a pilot, but those responsibilities should exist before work arrives.

We estimate where the manual process stops fitting. Half an hour of setup for five companies takes two and a half hours. For a hundred it takes fifty hours, competing with development time. This is not a sales forecast. It tells us what growth would require a different onboarding process or staffing choice.

Some assistance is part of learning, but repeated dependence is different. Explaining a status once and fixing it for someone every day are not the same activity. Pilot notes distinguish what users learn from what they cannot do independently. This helps the next version address the cause rather than simply hide the manual cost.

If assistance touches customer data, it follows a controlled process with limited access. Copying production records into personal spreadsheets for a convenient import creates consequences even in a small pilot. Simple implementation still needs ownership and a way to detect and correct mistakes. Small scale does not make those decisions disappear.

We automate work that has become frequent, understood and expensive. An occasional import whose format changes after every interview is a poor candidate for a universal builder. Daily notifications following a stable rule are easier to justify. The next release then responds to observed work rather than fear of hypothetical future scale.

Protect the conditions that make the experiment valid

If jobs disappear or one company can see another company's data, the experiment breaks down. Dispatchers keep a backup spreadsheet because they do not trust the service. Low adoption then tells us about reliability, not the value of the workflow. Backups, access controls and recovery need to support the work we ask people to put into the product.

The necessary quality floor depends on the promise. Manual account recovery may work for invited pilot users, while open signup requires another approach. Agree on supported conditions, participant expectations and failure recovery before people entrust real work to the service.

For jobs, reliability starts with specific questions. What happens after pressing create twice? Can one company open another company's record by changing a link? Can an accidentally deleted job be recovered? These checks protect promises made to participants rather than expand a generic best-practice checklist.

Perfect availability is unnecessary for every pilot. We can agree on support hours, volume limits and manual recovery. But recovery must be something the team can perform. "We have a backup" establishes little without checking access, completeness and the time needed to restore the job people are trying to finish.

We separate visual polish from workflow quality. A plain form can solve the problem well. A form with unclear validation forces employees to guess whether data was saved. That changes adoption. Legible states, confirmation and an understandable failure path can matter more to the experiment than another visual refinement.

Someone owns pilot incidents. A participant knows where to report a missing notification, and the team knows who checks delivery and helps work continue. Without ownership, every failure turns into a search for whichever developer happens to remember the relevant part of the system. The pilot then measures access to that person.

We also separate observation periods before and after a defect. If notifications failed in week one, low return rates cannot be attributed entirely to the idea's value. We record the affected period and recheck the workflow. This avoids both abandoning a useful product prematurely and excusing every weak result as an implementation issue.

Make scope changes an explicit decision

MVP scope usually grows one reasonable request at a time. A short brief keeps each discussion tied to the pilot: the question, the workflow, essential conditions, deferred features and release criteria. For a new request, we ask what we would be unable to learn without it.

Technician reporting can wait when the experiment tests job management. Correcting an accidentally assigned technician cannot: ordinary mistakes would stop the workflow. If a requirement is essential, name its cost and exchange another task, change the deadline or adjust the budget. Silently adding it breaks the scope agreement.

The brief describes behavior, not only feature names. "Notifications" is broad; "email the technician after assignment, without read receipts or SMS" can be estimated. A shared description of this version prevents more disagreement than a long list of headings whose meaning each participant interprets differently.

Deferred work gets a reason and a return signal. Bulk assignment might become useful when individual assignment consumes a significant part of a shift. Without a signal, the later list turns into promises nobody made. Reaching the second release date does not establish that every deferred feature is now necessary.

A scope discussion brings consequences together: which assumption the request affects, what work it adds, what leaves the plan and who decides. A developer should not guess business priority alone. A manager should not discover the new deadline at the end of the iteration after several small requests have accumulated.

We leave room for unknowns in estimates. An external integration may be harder than its documentation suggests. A short access and data check can precede a firmer estimate. That step buys information before committing to the full implementation, instead of quietly treating uncertainty as a precise delivery promise.

Scope can change when the pilot question changes. If companies already manage jobs well and only customer status causes trouble, protecting the original feature set no longer helps. We explicitly choose a different experiment and acknowledge the old one no longer answers the business problem. Scope control should preserve purpose, not an outdated list.

Use the pilot to decide what happens next

Review completed work and returning users, then investigate reasons. A company may stop because a status is confusing, demand is seasonal or another employee controls the buying decision. The completion rate cannot explain these differences by itself.

The pilot can justify continuing the workflow, changing the audience or dropping the idea. Willingness to pay needs its own check; a free pilot does not establish a price. The next release should address what we still do not understand. The deferred backlog does not become mandatory just because we shipped.

We review facts for each business: when it connected, jobs entered and finished, assistance required and recent active users. This prevents an attractive average from hiding one enthusiastic company and four abandoned accounts. The detail is useful because it points to different causes, not because a pilot needs another dashboard.

Interviews focus on recent actions. "Why did you stop?" may invite a polite general response. "Show us the last job you finished in the spreadsheet" brings the conversation back to work. A missing field, an ownership mismatch and a purchasing decision outside the pilot need different changes.

A price check uses a concrete offer: what the business pays for, support included and when payment begins. Agreement that someone would buy a useful service is cheap evidence. A real next step, such as continuing on stated paid terms, is closer to testing the business model than an abstract expression of interest.

We connect the next decision to its cost. One more week correcting the workflow and six months building a mobile app need different evidence. The more expensive the commitment, the stronger the grounds should be. Avoiding a large investment on weak information can be a good outcome of an MVP.

Finally, we preserve findings with limits: participants invited, assistance provided, missing measurements and remaining questions. Those details fade quickly, while "the pilot succeeded" starts justifying unrelated decisions. A short honest record lets the next version investigate new uncertainty rather than treat the old uncertainty as a settled fact.

Putting the pilot plan together

Our repair-company pilot asks whether dispatchers can manage jobs without the parallel spreadsheet. Creation, assignment, rescheduling, status and completion stay in the first release. One repair category and one role reduce variations without breaking the journey. Technicians respond through their familiar channel; customers receive a status page.

Before invitations we test a normal job, wrong assignment, a moved appointment and duplicate creation. Company separation, recovery and clear failures protect the experiment. Setup follows a manual instruction, and pilot notes record assistance separately from independent use. Otherwise work completed by the team can masquerade as customer adoption.

For each company we capture the starting process and weekly results. After four weeks we review recent jobs, return behavior and a concrete paid continuation. The time box limits this experiment; it does not prove every commercial assumption. Contradictory evidence calls for a small follow-up that separates causes.

New features face the same question. A calendar belongs when its absence genuinely blocks assignment and sends people back to the spreadsheet. Analytics waits while it is irrelevant to the current decision. The pilot produces facts, reasons and a next check, giving the team a stronger basis before a larger investment.

Back to articles