How this study works — and what it's for.
This page documents the research behind the Artisan survey: what we are trying to learn, how every question earns its place, how the pricing modules work, the decision thresholds we committed to before collecting a single answer, and how responses are handled. It exists because research you can't inspect is just marketing.
Taking the survey? Ideally answer first and read this after — not because anything here is secret, but because unprimed answers are better data.
Artisan is a marketplace that makes construction-labour reliability visible, portable and rewarded — starting with steel fixing in Sydney. Before scaling it, four founder assumptions must be replaced with first-party evidence:
- That short-notice hiring fails often enough, and expensively enough, to change behaviour — measured in no-shows per month and dollars per failure.
- That the informal system (calls, WhatsApp, favours) is where hiring actually happens — measured by recalled last-time behaviour, not opinion.
- That bosses will pay for guaranteed, reliability-rated labour — measured by two independent pricing methods plus commitment signals.
- That workers want reputation that travels and faster pay — measured by stated appetite and the willingness to join a launch list.
Every number this study produces replaces an estimate in the Artisan business plan with a figure that carries an N. That is its entire job.
Two structured questionnaires — bosses/hirers (6 steps, ~7 min) and workers/crews (6 steps, ~5 min) — plus follow-up interviews with story-leavers.
Construction trade labour, beachheaded on the Mongolian steel-fixing community in Sydney; small external control group to test generalisation.
Purposive (named, known respondents first) then snowball via referrals — appropriate for a trust-dense, hard-to-reach trade population; not a probability sample and not presented as one.
Self-completed on mobile, distributed person-to-person with per-channel source tags. Fully bilingual — every question, option and screen has a natural Mongolian translation behind a one-tap EN/МН toggle, and each response records the language it was answered in.
Past behaviour over hypotheticals, specifics over generalities, stories invited, money quantified — and commitment signals (contact, pilot, referrals) treated as the highest-grade evidence.
Live aggregation by segment on the admin dashboard: distributions for categorical answers, medians/quartiles for numerics, PSM intersection charts for pricing, CSV export for deeper work.
Sizes the respondent: role (owner vs foreman — who actually buys), crew on books, weekly labour need, active sites and the current day-rate. Trade is assumed — steel fixing, the beachhead — so the question was removed to keep the survey short. This anchors every later answer and builds the first-party day-rate dataset that replaces our published estimates.
“Last time you needed a worker at short notice — how did you find them?” is deliberately past-tense and specific. Recalled behaviour resists flattery; hypotheticals don't. Fill frequency and time-to-fill quantify how often the problem occurs and how long it bleeds.
No-show counts, the cost band of a lost day, slipped pours and wrong-hire rework. This is the section the business case rests on: it converts anecdote into a distribution of measured losses — and the free-text story fields capture the quotes that survive into the business plan.
What makes a boss say yes to a stranger today (vouching, seen work, tickets, price, availability), who actually decides — and the sharpest question in the study: would they pay a per-day premium for a worker with a proven reliability record? That answer prices the product's core promise directly.
Van Westendorp price sensitivity plus a Gabor-Granger ladder, with a take-rate alternative as a cross-check. Two independent methods that converge give a defensible price; one method alone gives a guess (see the pricing science below).
Optional contact details, a yes/no on a free pilot fill, and a referral count. Talk is cheap — a phone number, a pilot acceptance and introductions are skin in the game. These three fields are the strongest evidence in the entire study.
Experience, employment mode (sole trader / crew / employed), tickets held and city — trade is assumed (steel fixing, the beachhead). Maps the supply side's shape and qualifies every answer that follows.
Last-job channel (past-tense again), idle days in the last quarter, lead-time on the next job, and the single biggest frustration — named in the worker's own categories. Idle days are the supply side's no-show: measurable waste the marketplace exists to remove.
Current day-rate, late payments in the last year, and whether they've ever simply not been paid. Non-payment stories justify the payment-protection roadmap and measure how broken trust is in the other direction.
What makes a worker choose a boss, and what proof of skill they can show today (usually: nothing portable). This is the demand for the product's core promise — reliability that travels.
An instant-pay Gabor ladder (percentage of wages for same-day payment) and appetite for a visible reliability profile. Tests the worker-side monetisation thesis without ever paywalling access to work.
Optional name and mobile plus a referral count — “first on the list for paid shifts” converts research goodwill directly into launch supply, and completion earns a Founding 100 spot (an honest, costless reward: founding status with real launch perks, recorded with the response).
Asking “what would you pay?” produces polite fiction. The study uses two established techniques and a cross-check, all anchored to one concrete product description rather than an abstract idea:
A vetted, reliability-rated fixer delivered to your site for tomorrow — guaranteed to show, or we instantly re-fill and credit you.
Four price points per respondent: too cheap to trust, a bargain, getting expensive, too expensive to consider. Plotted across the sample, the curves intersect in an acceptable price range and an optimal price point. Designed for exactly this situation — pricing a product category the market hasn't seen before.
A fixed ladder of monthly price points (A$49 → A$299+); each respondent marks the highest they'd accept. This yields a demand curve and the revenue-maximising price — and, run against the PSM result, a convergence test: if the two methods disagree wildly, neither is trusted.
The same value reframed as a commission on a A$400 day instead of a subscription. Some buyers reject subscriptions but accept transaction pricing (or vice versa) — this question decides which revenue model the market prefers, not just the number.
Workers get a parallel module: an instant-pay ladder priced in percent of wages — testing supply-side monetisation that never paywalls access to work.
The point of committing to numbers in advance is that the evidence can change our minds — and enthusiasm can't. Each signal below has a go threshold and a kill threshold; results in between trigger deeper interviews, not a coin flip.
| Signal | Go | Kill |
|---|---|---|
| Problem raised unprompted in follow-up interviews | ≥ 60% of bosses | < 30% |
| Short-notice fills are frequent AND costed | ≥ 2/month with $ attached | Rare or uncosted |
| Boss willingness to pay (PSM/Gabor midpoint) | Supports ≥ A$99/mo | Clusters at A$0–49 |
| Free-pilot acceptance (bosses) | ≥ 40% say yes | < 15% |
| Referrals volunteered per respondent | ≥ 1.5 average | < 0.5 |
| Workers open to a visible reliability profile | ≥ 50% yes/maybe | Majority no |
Targets: 30–50 bosses/foremen and 40–60 workers, starting inside the Mongolian steel-fixing network in Sydney where trust is pre-existing, then outward through referrals. A small non-Mongolian control group (5–10) tests whether findings generalise.
Every link you share should carry a source tag — /s/boss?src=wa-group-1, ?src=event-jun, ?src=referral. Source is stored on each response, so channel quality (completion rate, contact rate, referral rate) is measurable per channel.
A personal message from a known name outperforms any broadcast. Lead with reciprocity (“you'll get the results back”) and the time cost (“5–7 minutes”), never with the product. The survey must read as research, not marketing — because it is.
Respondents who leave a pour-slip story, a non-payment story or a phone number get a 15–20 minute call. The quantitative survey sizes the problem; the calls supply the understanding (and the quotes). Ask about the past, never about the future.
Aggregates update live on the admin dashboard. Interim peeks are fine for fieldwork steering, but go/kill decisions wait until the target sample is reached — the thresholds were set in advance precisely so enthusiasm can't move them.
- Responses are anonymous by default — contact details are optional, stored in a separate field, and used only for the stated purpose (results share-back, pilot list).
- Results are reported in aggregate only; free-text stories are quoted without identifying details.
- Each response stores its answers, segment, source tag and completion metadata (duration, steps) in Supabase with row-level security; reading results requires a passphrase-protected RPC — there is no public read path.
- Completion time is used as a quality filter: sub-90-second submissions are flagged as speeders during analysis.
- The schema is versioned by stable question ids — questions can be added without breaking historical aggregation, and changed ids are treated as new questions.
This is founder-led research on a purposive sample inside a community the founder belongs to. That is its strength — access and honesty money can't buy — and its bias: respondents may be kinder than strangers, and the beachhead may not represent the wider market. The control group, the pre-registered thresholds, the preference for behaviour over opinion and the weight placed on costly commitment signals all exist to counter exactly that. The findings will be read as evidence about the beachhead first, and generalised with care.