A six-hour request hides a production system
A take-home advertised as six hours rarely takes six. The clock covers the visible artifact, not decoding an unfamiliar business, choosing tools, cleaning supplied material, deciding what “good” means, and second-guessing an unstated audience. Candidates with spare evenings, fast laptops, and category exposure start differently from equally capable people without them. A team may think it is testing job skill while rewarding tolerance for ambiguity and unpaid availability. Cutting the duration to three hours is not simply requesting fewer slides. It requires a smaller unit of evidence: one decision, bounded inputs, and review that values judgment over production polish.
The hidden work behind each task
A useful work sample has five parts: a decision, its constraints, available evidence, an artifact recording reasoning, and a scoring rubric. Omit one and candidates must invent it. “Audit our onboarding and propose improvements” silently requests market research, product interpretation, prioritisation, copywriting, design taste, and assumptions about engineering capacity. The final deck may be short; the invisible work is not.
Estimate time from the least informed qualified person, not the manager who wrote the brief. Pilot it with a colleague lacking internal context, and record setup, reading, research, analysis, writing, and final checks separately. Staff have product knowledge, shared vocabulary, and people who can resolve vague phrases in minutes; candidates do not. If a reviewer must explain the intended answer after reading a submission, the prompt lacked context.
A time cap differs from fixed output. “Spend no more than three hours” alongside a polished strategy, model, and presentation still pressures conscientious candidates to exceed the cap. A credible three-hour brief states expected effort and the maximum acceptable deliverable, including what may remain unfinished. That lets candidates show prioritisation rather than hide compromises required in a real role.
Three scoping models and their trade-offs
The diagnostic memo is the lowest-friction model. Give candidates a constrained scenario, short data pack, and decision owner, then request a one- or two-page recommendation. A product candidate might choose which of three activation problems to investigate; a marketing candidate might select a channel experiment and define its success measure. This reveals framing, prioritisation, and clarity while placing little weight on presentation software or paid research tools. Its weakness is limited evidence of craft where craft is the work.
A bounded artifact tests craft more directly. A designer can improve one screen from a supplied flow, an analyst answer one question from a clean dataset, and a content strategist revise a short landing-page section for a stated audience and objective. Guardrails must be firm: provide source files, declare intended fidelity, and limit states or outputs. Otherwise it becomes a contest in unpaid production time. It can also resemble work the company could ship, creating ethical and legal issues.
A case review with a structured conversation moves part of assessment into paid interview time. Candidates receive a concise case, write notes or a short response, then explain choices and respond to a changed condition in a 30- to 45-minute session. It reveals thinking under challenge without requiring a glossy deliverable and suits senior, client-facing, and cross-functional roles. It does require trained interviewers; unstructured discussion can replace one bias with another.
No model is universally right. Use a diagnostic memo when judgment is the hiring question, a bounded artifact for observable technical or creative output, and a live case where communication and adaptation matter. Do not ask one exercise to prove every capability. Distribute evidence across portfolio review, structured questions, a short work sample, and reference checks.
Where candidate effort gets wasted
Output-based estimates create bad scope. A manager sees ten slides and assumes two hours; a candidate sees interpreting the brief, finding evidence, deciding the story, creating charts, formatting the deck, and preparing for aesthetic judgment. Combining a market scan, strategy proposal, financial model, and executive presentation does not test range; it creates four tasks with different failure modes.
Ambiguity is sometimes defended as initiative. In a real job, initiative includes asking questions, finding owners, and making assumptions visible. A take-home with no stakeholder access tests none of this; it rewards guessing the interviewer’s private preference. Use a shared question log for a fixed window and publish every answer to all participants. If questions are prohibited, state the assumptions to use instead of making uncertainty a trap.
Tool expectations create another divide. A specialist design suite, paid data source, powerful machine, or proprietary AI subscription can exclude people before review. State optional tools, sufficient source material, and whether AI assistance is permitted. A blanket ban is difficult to assess and rarely matches workplace reality; unrestricted, undisclosed use can make authorship impossible to judge. Instead, ask candidates to name tools used, describe material they generated or edited, and defend choices in conversation.
Build a three-hour assignment around one decision
Start with the hiring decision, not the deliverable. Finish: “After reviewing this work, we will know whether the candidate can…” Describe a role-relevant behaviour, such as prioritising an ambiguous problem, translating data into a recommendation, or producing clear interface copy under constraints. If it contains three verbs joined by “and,” split the evidence across the interview process.
A practical 180-minute envelope:
- 15 minutes — read the brief and success criteria. Include role context, audience, decision owner, constraints, source material, output format, and allowed tools.
- 35 minutes — inspect a limited evidence pack. Supply relevant metrics, customer excerpts, design files, or code fragments; do not make source gathering the assignment.
- 65 minutes — make and support a decision. Request priorities, assumptions, trade-offs, and evidence that would change the recommendation.
- 45 minutes — create a compact artifact. Cap it at a two-page memo, one analysis notebook, three screens, or an equivalent discipline-appropriate format.
- 20 minutes — check and annotate. Let candidates flag unfinished areas and say where they would spend their next hour.
This final step is more revealing than extra polish. People who identify risk, distinguish evidence from assumption, and name the next validation step make work legible to collaborators. Do not demand exact time tracking as proof of honesty; the cap is an employer design constraint. If a pilot participant cannot finish without cutting meaningful work, reduce the brief before sending it to candidates.
Pilot with someone who understands the role but lacks insider knowledge of the scenario. Ask where they paused, what they assumed, and which material they ignored. Remove anything that does not feed the rubric. A shared question log can remain open for two business days without extending the effort cap if all candidates receive the same answers and the deadline moves only when the company changes the brief.
Can you assess the work fairly?
A shorter task earns trust only with disciplined review. Build the rubric before inviting candidates and score independently before panel discussion. Four criteria are usually enough: problem framing, use of evidence, quality of trade-offs, and communication for the stated audience. Give each observable language. “Demonstrates strategic thinking” invites preference; “states a priority, names what is deferred, and ties the choice to supplied evidence” tells reviewers what to find.
Do not grade internal vocabulary, unsupplied knowledge, or polish associated with an agency production budget. A clean submission is easier to review, but visual finish should not outweigh the tested ability unless final-form visual craft is the role. Reviewers need the same materials and calibration examples before scoring. Without calibration, a structured rubric is labels attached to instinct.
Address the assessment/free-labor boundary directly. If the company could use a submission in production, use fictional, expired, or heavily altered material, or pay for the project. Candidates should not sign broad intellectual-property terms for a short screening exercise. Give clear confidentiality instructions, avoid sensitive customer data, and offer an accessible alternative when a standard format creates a barrier. Fairness is built into workload, materials, review, and rights, not a disclaimer.
Hiring signals are changing with AI
Long take-homes grew from the belief that more output gives better evidence. AI can now generate plausible decks, copy, code, and research summaries in minutes, weakening that belief. The answer is not a larger artifact or policing every tool, but assessing ownership: how candidates framed the problem, checked source quality, selected options, and identified output limits.
Ask for an assumption register, a short note on tools used, or a critique of the candidate’s own draft. Follow-up questions can test understanding: Which metric would you distrust first? What would reverse the priority? Which output would you validate with a customer or stakeholder? Generated polish matters little when reviewers can see reasoning and probe it consistently.
The stronger design is not asynchronous versus live, or manual versus AI-assisted work. It is comparable evidence with proportionate burden. A concise work sample, structured interview, and calibrated review panel can reveal more day-to-day judgment than a weekend assignment only the most available candidates can complete.
A smaller brief produces clearer evidence
Three hours is enough when an assignment asks one meaningful question and supplies the conditions to answer it. Reduce research the company can provide, cut deliverables that do not map to the rubric, and preserve room to show assumptions and trade-offs. The submission may look less finished than a six-hour deck, but it is easier to compare and review and less likely to confuse endurance with ability. That is a better hiring signal for both sides.