How We Ship Custom AI Systems in 6 Weeks (and Why It Works)

The six-week engagement model in detail — what happens each week, the four rules that make the timeline real, and the conditions under which it does not work.

Six weeks sounds like a sales number. It is not — it is a constraint we impose on ourselves because it forces a set of decisions that make projects succeed, and removing it reliably makes them fail. The constraint is doing the work.

To be precise about the claim: six weeks is the time to a first deployable — a system running in your environment, on your data, measured against a target metric agreed in week one. It is not the time to a finished roadmap, and it is not a promise that one use case is all you will ever want. It is a promise that you will know whether this works, with working software in hand, before a quarter has closed.

Why the deadline is the design tool

Give an AI project six months and something specific happens. The scope grows to fill the time, because every stakeholder who hears about it has an adjacent request and no reason to say no. The team explores architectures instead of choosing one. Evaluation gets deferred because the model is “not ready to measure yet.” And then at month five, the project meets reality — real data, real users, real integration — with five months of accumulated assumptions that have never been tested.

Six weeks makes that impossible. You cannot gold-plate, you cannot defer evaluation to the end, and you cannot avoid talking to the people who will use the thing. The timeline does the prioritisation that committees are bad at.

Week 1 — Discover

The entire week is spent not building, which clients find uncomfortable and which is where most of the project’s eventual success is determined.

We interview the people who do the work today, not only the people who sponsor the project. Those are different populations with different information, and the operators know where the real exceptions live.

Then we measure the baseline, and this is the step most projects skip. Almost nobody knows their current performance with any precision. They have a belief — “inspection catches most defects,” “review takes about half an hour” — and the belief is usually wrong in an interesting direction. On a vision inspection project, two inspectors independently grading the same 800 parts agreed with each other 86% of the time. That single number reframed the entire project: the target was no longer perfection, it was beating a measurable 86% at twenty-five times the coverage.

Week one ends with an opportunity brief: baseline, target metric, data inventory, and a go/no-go recommendation. Roughly one engagement in five ends here, with us recommending that you not build. That is the cheapest possible good outcome and the reason discovery is priced and contracted separately from the build.

Week 2 — Architect

One week to make the decisions that are expensive to reverse, written down before any production code exists.

Model and retrieval strategy. Deployment topology — and this is where constraints surface that would otherwise ambush the project in week five. “No outbound internet from the production VLAN” changes everything about model selection; discovering it late has killed more AI projects than any modelling mistake.

Also in week two, two things that are easy to postpone and should not be:

The evaluation harness is designed now, before the model exists. If you cannot describe how you will measure success, you are not ready to build — and the act of specifying the measurement usually exposes an underspecified requirement.

Cost and latency are modelled now. Tokens per request times requests per day times price, or GPU hours at your utilisation. Unit economics discovered after a build are a bad conversation; we would rather have it in week two when the architecture can still change.

Week two produces a technical architecture, an evaluation plan and a fixed-scope build estimate. We quote fixed scope because by this point we know enough to be accountable for it.

Weeks 3–4 — Build

Two weeks of building, structured around weekly demos against real data. Not slides — the actual system, in its actual state, including what is broken.

The sequencing rule is end-to-end before deep. By the end of week three something works badly all the way through: data in, model inference, output in the real interface. A thin complete path beats three excellent components with no path between them, because it surfaces integration problems while there is still time, and because the ugly end-to-end version is what makes the stakeholder conversation concrete.

Week four is depth: improving the components the evaluation harness says are weakest. The harness built in week two is what makes this an engineering activity rather than an argument about which model feels better.

Week 5 — Validate

A full week on proving it works, which is not a step that can be compressed into the build.

Offline evaluation against the held-out gold set, segmented — not one aggregate number, but accuracy by category, by difficulty, by data source. Aggregates hide the failure modes that matter, and the segment where the system is weak is usually the segment the client cares about most.

Then shadow mode or A/B against the current process. The model runs, logs everything, and affects nothing. Every disagreement between system and human gets adjudicated by a domain expert against ground truth. That adjudication set is the real deliverable of week five — it is what moves the operating threshold from a guess to a decision.

Then red-teaming: adversarial inputs, bias review, failure-mode analysis. What does it do with an input it has never seen? With an input designed to mislead it? With a decoy document that looks correct?

Week five ends with a pass/fail against the week-one target metric. Explicit, written, no reinterpreting the goal after seeing the results.

Week 6 — Deploy and scale

Production rollout with monitoring, alerting and a rollback path that has been tested rather than assumed. Runbooks and dashboards. Training for the team that will own it.

Then handover: source code, weights, evaluation datasets, infrastructure-as-code, documentation, all in your repositories. No runtime dependency on us unless you specifically ask for one.

The four rules that make it real

1. One use case. Fixed. The scope is one problem with one success metric. Adjacent requests go on a roadmap for a later engagement, not into this one. This is the rule clients push hardest on and the one that most determines whether six weeks happens.

2. Data access exists before week three. Not “is being arranged.” Exists, tested, with a real credential. The most common cause of slippage in this industry is not technical difficulty, it is a provisioning ticket. We will not start a build phase on a promise of access.

3. One decision-maker, available weekly. A named person who can approve a tradeoff in a day. Design by committee cannot fit in six weeks; a weekly 45-minute decision meeting can.

4. Evaluation before optimisation. No model improvement work until the harness exists. Otherwise “better” is a matter of opinion and the team optimises whatever is most interesting.

When it does not work

Honesty is more useful than a universal claim. Six weeks is the wrong model when:

  • The data genuinely does not exist. If you need to instrument a process to generate training data, that is a separate and worthwhile project first. We will say so in week one.
  • The problem is organisational. If the real issue is that two departments disagree about who owns a decision, no model fixes that, and shipping one makes the disagreement more expensive.
  • The use case is irreducibly broad. “Put AI across customer service” is a programme, not an engagement. We will help you decompose it, then do the first piece in six weeks.
  • Regulatory approval gates the deployment. We can build and validate in six weeks; we cannot compress your regulator’s review cycle, and we will not pretend to.

Why clients end up preferring the constraint

The reaction in week one is usually scepticism. The reaction in week three — when something ugly but real is running on their own data — is where it changes. Not because the system is good yet, but because the uncertainty collapsed. They can see the shape of the thing, argue about specifics with evidence, and decide whether to keep going with information instead of faith.

Six months of confidence is worth less than three weeks of evidence. That is the whole argument.


See the full engagement model with deliverables per phase, or book a consultation to scope a first use case.

Let's scope your AI opportunity

A 45-minute conversation is usually enough to tell whether there is a system worth building — and what it would take. No obligation, no deck.