How we work

Artifacts over advice. Evidence over confidence.

Two commitments shape everything below. Every engagement produces machine-readable artifacts that run in your systems without us. And every claim we make — including the ones on this website — carries a grade, so you can tell the difference between a measured result and a vendor’s press release.

Four bounded delivery stages separated by evidence gates and a red boundary that prevents skipping ahead.

Forward-deployed, in two-week increments

We work inside your estate, on your systems, alongside your engineers — not in a parallel universe that gets handed over at the end. The unit of progress is a working artifact in your repository, not a status update.

Two-week increments

Every fortnight produces something that runs. If a fortnight produces only findings, we say so plainly and change what we are doing rather than reporting progress toward a milestone nobody can inspect.

Fixed fee per bounded outcome

Never hourly. New priorities replace old ones or take a change order. Our delivery efficiency is our margin problem, not a discount we owe you and not a reason to inflate scope.

Capability transfer is the goal

We are trying to make ourselves unnecessary for the rung you bought, so you can decide freely whether to buy the next one. Dependency is a business model we are not interested in.

The module acceptance contract

Six conditions, agreed before work starts, applied to every module we deliver. The last one is the one that matters.

  • The deliverable is machine-readable and lives in your repositories and CI
  • It has a named owner on your side, agreed at the start rather than found at the end
  • Its acceptance evidence is stated in advance and is objectively checkable
  • Failure behaviour is documented and demonstrated, not just described
  • Rollback and disable paths are tested in front of you
  • Your operator can run the documented procedure without the implementation team in the room

Every workflow has a stated ceiling and a written route to raise it

We do not deliver “an agent”. We deliver a workflow at a named autonomy level, with the evidence required to advance to the next one written down before anyone starts. No workflow advances because a model claims confidence — advancement is an accountable business and risk decision supported by measurement.

LevelAuthorityEvidence to advance
0 — ObserveRead approved data, produce private evidenceAccess and provenance tests
1 — RecommendPropose an action; a person performs itQuality, usefulness, recorded rejection reasons
2 — PreparePrepare an exact change for approvalDry-run, validation, rollback, approval binding
3 — Execute boundedReversible, low-consequence actions inside explicit limitsSustained success, monitoring, recovery, low override rate
4 — Execute with escalationOperate a bounded workflow, escalate deviationsMature control evidence, named operational ownership

The same discipline applies to our own delivery. We use agents heavily to do this work — that is part of how a small team delivers at this scope — and they operate under the same rules we sell you. They cannot approve their own output, grant themselves access to your systems, or bypass a human on a consequential action. Every artifact we hand over has a human who read it and is accountable for it.

— AI-assisted delivery under human control

What you keep, in writing

Yours from day one

  • All code, infrastructure definitions, contracts, schemas, and policies
  • The golden task suite and every evaluation we build with you
  • The context graph, the agent registry, and the outcome ledger definitions
  • Runbooks, drill records, incident procedures, and decision logs
  • A stated exit path, tested rather than promised

What we keep

Reusable product components — the graph builder, the provenance and evaluation harness, the state schema, the tool contract kit, the ledger — which we improve across engagements and deploy into yours. You get the benefit of every previous client’s hard lessons; nobody gets your data, your graph, or your task suite.

If a component is genuinely bespoke to you, it is yours outright. The 70/20/10 rule keeps that boundary honest rather than convenient.

We grade our sources, and we publish the ones that weaken our case

This is a market where most public material is produced by people selling into it. Our response is a grading scheme applied to every figure we use — in a proposal, in a report, or on this website.

Grade A

Primary source

A published study, benchmark, or dataset we can read directly and whose method we can inspect. Quotable without qualification.

Grade B

Credible secondary

Reputable reporting on a primary source we have not read in full. Usable, always attributed, never load-bearing on its own.

Grade C

Single or self-interested

One source, a vendor benchmark, or an unnamed survey. Directional only. Never priced against, and re-tested on your data before it informs a decision.

Things we will not oversell

Regulatory urgency

The EU AI Act’s high-risk obligations were pushed to December 2027. Anyone using regulatory panic to compress your decision timeline is selling, not advising. Build the controls because they make agents work, not because of a deadline that moved.

Incident statistics

Reported rates of agent-linked security incidents in enterprises range widely across sources, and several of the most quoted figures come from vendors with a product to sell. The direction is real; the specific numbers are not something we will price your engagement against.

Semantic layer benchmarks

The 90% → 98% figures are published by semantic-layer vendors. That is exactly why our deliverable is the same test run on your twenty real business questions rather than a slide quoting theirs.

Our own market corpus

Our internal signal analysis draws on hundreds of thousands of words of practitioner and vendor commentary. It is attention signal, not demand evidence — heavy on founders and tool creators, light on buyers and failed adopters. It sharpens our failure modes. It cannot establish that anyone will pay.

Selected sources

The objections, answered as we would answer them in the room

Why would we not just build this ourselves?

Often you should, and if that is the right answer we will tell you which parts. You have the platform, the context, and the credibility internally — those are real advantages.

Two arguments against doing all of it yourself. First, the evidence: reported pilot success rates for teams blending internal specialists with external expertise are substantially higher than internal-only builds (67% versus 22% in MIT’s coverage, grade B). Second, the specific failure modes here — false completion, lifecycle escape, resumption failure, silent data defects — are ones you learn by hitting them repeatedly across many estates, and each one is expensive to learn on your own production systems.

The honest split: you build what encodes your context. We bring the harnesses, the failure catalogue, and the controls that are the same everywhere.

Why not just buy an agent observability or guardrails product?

You probably should buy one, and we will help you evaluate them. Guardrails and observability are the most heavily funded corridor in this market and the products are genuinely good at what they do.

What they do not do is catch a plausible artifact sourced from the wrong place, hold durable execution state across sessions, encode your estate’s reality into a graph an agent can use, or prove that your kill switch survives a restart. A product ships a generic capability. The defensible part is account-specific, and it is not going to arrive in a release note.

What if a hyperscaler ships all of this as a platform feature?

Parts of it will be commoditised, and we plan for that rather than against it — we build for a roughly 24–36 month defensibility horizon, not a ten-year one. Generic versions of the context graph, the tool plane, and evaluation harnesses will ship from vendors.

What does not ship in a product release: the graph of your estate, your golden task suite with your traps in it, and the operating relationship that lets someone hold production and data permissions responsibly. If a vendor collapses part of our offer into a default feature, the correct response is to adopt it and move up the stack. We would rather say that now than defend a moat that has already been drained.

So are you a consultancy or a product company?

Services-funded product. Every engagement is contractually shaped to leave a reusable component behind, and we cap engagements that produce none. That is deliberately a different thing from a consultancy that hopes to become a product company later, and it is also not a startup burning capital to build a platform nobody has asked for yet.

The practical consequence for you: we are motivated to make delivery repeatable rather than to maximise billable hours, and we will decline work that would fork our core into something unmaintainable.

How mature is this? Do you have reference clients?

We are early, and we would rather you hear that from us. The model, the failure catalogue, the architecture, and the evidence base are real and documented in depth. Rungs 1–3 are capability we can deliver now. Rungs 4 and 5 are a published destination we would deliver with partners or a larger team, and we do not imply otherwise.

Price bands on this site are working ranges anchored to market benchmarks, not to a long history of closed deals. If you need a vendor with fifty logos and a procurement-friendly reference list, we are not that yet, and pretending otherwise would waste your time. What you get instead is senior attention, an unusually explicit method, and a company for whom your outcome is existential.

What happens if we stop working with you?

Everything keeps running, because that is an acceptance condition rather than a courtesy. The artifacts are in your repositories, the procedures are documented, your operator has already run them unaided, and the exit path was tested rather than promised.

We are deliberately built so that leaving is cheap. A relationship that survives because exit is expensive is not one we want.

Where do you draw the line on what agents may do?

Money, access grants, legal commitments, regulated decisions, and destructive operations are never delegated by default. An AI system cannot approve its own release, grant itself production access, or bypass external authorization for a consequential action.

In a first engagement we also avoid anything touching customers, money, or regulated decisions entirely — the first workflow is engineering-facing with an internal blast radius, because that is where you learn cheaply.

Our platform is not ready. Should we still talk?

Yes, but expect us to say so. If you have no CI/CD, no infrastructure as code, and no cloud governance, there is nothing to install an agent runtime onto, and selling you one would be malpractice with an invoice attached.

That platform work is real work we know how to do, and we will scope it as its own engagement with its own shape. What we will not do is disguise it as an AI programme because that is what has budget this year.

Ask us the question you think we will dodge.

We have published the failure modes, the graded evidence, the price bands, the exclusions, and the limits of our own capability. The first conversation is 45 minutes and one question: where are you on the ladder, and what has already gone wrong?