Two-week increments
Every fortnight produces something that runs. If a fortnight produces only findings, we say so plainly and change what we are doing rather than reporting progress toward a milestone nobody can inspect.
Two commitments shape everything below. Every engagement produces machine-readable artifacts that run in your systems without us. And every claim we make — including the ones on this website — carries a grade, so you can tell the difference between a measured result and a vendor’s press release.
We work inside your estate, on your systems, alongside your engineers — not in a parallel universe that gets handed over at the end. The unit of progress is a working artifact in your repository, not a status update.
Every fortnight produces something that runs. If a fortnight produces only findings, we say so plainly and change what we are doing rather than reporting progress toward a milestone nobody can inspect.
Never hourly. New priorities replace old ones or take a change order. Our delivery efficiency is our margin problem, not a discount we owe you and not a reason to inflate scope.
We are trying to make ourselves unnecessary for the rung you bought, so you can decide freely whether to buy the next one. Dependency is a business model we are not interested in.
Six conditions, agreed before work starts, applied to every module we deliver. The last one is the one that matters.
We do not deliver “an agent”. We deliver a workflow at a named autonomy level, with the evidence required to advance to the next one written down before anyone starts. No workflow advances because a model claims confidence — advancement is an accountable business and risk decision supported by measurement.
| Level | Authority | Evidence to advance |
|---|---|---|
| 0 — Observe | Read approved data, produce private evidence | Access and provenance tests |
| 1 — Recommend | Propose an action; a person performs it | Quality, usefulness, recorded rejection reasons |
| 2 — Prepare | Prepare an exact change for approval | Dry-run, validation, rollback, approval binding |
| 3 — Execute bounded | Reversible, low-consequence actions inside explicit limits | Sustained success, monitoring, recovery, low override rate |
| 4 — Execute with escalation | Operate a bounded workflow, escalate deviations | Mature control evidence, named operational ownership |
The same discipline applies to our own delivery. We use agents heavily to do this work — that is part of how a small team delivers at this scope — and they operate under the same rules we sell you. They cannot approve their own output, grant themselves access to your systems, or bypass a human on a consequential action. Every artifact we hand over has a human who read it and is accountable for it.
— AI-assisted delivery under human controlReusable product components — the graph builder, the provenance and evaluation harness, the state schema, the tool contract kit, the ledger — which we improve across engagements and deploy into yours. You get the benefit of every previous client’s hard lessons; nobody gets your data, your graph, or your task suite.
If a component is genuinely bespoke to you, it is yours outright. The 70/20/10 rule keeps that boundary honest rather than convenient.
This is a market where most public material is produced by people selling into it. Our response is a grading scheme applied to every figure we use — in a proposal, in a report, or on this website.
A published study, benchmark, or dataset we can read directly and whose method we can inspect. Quotable without qualification.
Reputable reporting on a primary source we have not read in full. Usable, always attributed, never load-bearing on its own.
One source, a vendor benchmark, or an unnamed survey. Directional only. Never priced against, and re-tested on your data before it informs a decision.
The EU AI Act’s high-risk obligations were pushed to December 2027. Anyone using regulatory panic to compress your decision timeline is selling, not advising. Build the controls because they make agents work, not because of a deadline that moved.
Reported rates of agent-linked security incidents in enterprises range widely across sources, and several of the most quoted figures come from vendors with a product to sell. The direction is real; the specific numbers are not something we will price your engagement against.
The 90% → 98% figures are published by semantic-layer vendors. That is exactly why our deliverable is the same test run on your twenty real business questions rather than a slide quoting theirs.
Our internal signal analysis draws on hundreds of thousands of words of practitioner and vendor commentary. It is attention signal, not demand evidence — heavy on founders and tool creators, light on buyers and failed adopters. It sharpens our failure modes. It cannot establish that anyone will pay.
Often you should, and if that is the right answer we will tell you which parts. You have the platform, the context, and the credibility internally — those are real advantages.
Two arguments against doing all of it yourself. First, the evidence: reported pilot success rates for teams blending internal specialists with external expertise are substantially higher than internal-only builds (67% versus 22% in MIT’s coverage, grade B). Second, the specific failure modes here — false completion, lifecycle escape, resumption failure, silent data defects — are ones you learn by hitting them repeatedly across many estates, and each one is expensive to learn on your own production systems.
The honest split: you build what encodes your context. We bring the harnesses, the failure catalogue, and the controls that are the same everywhere.
You probably should buy one, and we will help you evaluate them. Guardrails and observability are the most heavily funded corridor in this market and the products are genuinely good at what they do.
What they do not do is catch a plausible artifact sourced from the wrong place, hold durable execution state across sessions, encode your estate’s reality into a graph an agent can use, or prove that your kill switch survives a restart. A product ships a generic capability. The defensible part is account-specific, and it is not going to arrive in a release note.
Parts of it will be commoditised, and we plan for that rather than against it — we build for a roughly 24–36 month defensibility horizon, not a ten-year one. Generic versions of the context graph, the tool plane, and evaluation harnesses will ship from vendors.
What does not ship in a product release: the graph of your estate, your golden task suite with your traps in it, and the operating relationship that lets someone hold production and data permissions responsibly. If a vendor collapses part of our offer into a default feature, the correct response is to adopt it and move up the stack. We would rather say that now than defend a moat that has already been drained.
Services-funded product. Every engagement is contractually shaped to leave a reusable component behind, and we cap engagements that produce none. That is deliberately a different thing from a consultancy that hopes to become a product company later, and it is also not a startup burning capital to build a platform nobody has asked for yet.
The practical consequence for you: we are motivated to make delivery repeatable rather than to maximise billable hours, and we will decline work that would fork our core into something unmaintainable.
We are early, and we would rather you hear that from us. The model, the failure catalogue, the architecture, and the evidence base are real and documented in depth. Rungs 1–3 are capability we can deliver now. Rungs 4 and 5 are a published destination we would deliver with partners or a larger team, and we do not imply otherwise.
Price bands on this site are working ranges anchored to market benchmarks, not to a long history of closed deals. If you need a vendor with fifty logos and a procurement-friendly reference list, we are not that yet, and pretending otherwise would waste your time. What you get instead is senior attention, an unusually explicit method, and a company for whom your outcome is existential.
Everything keeps running, because that is an acceptance condition rather than a courtesy. The artifacts are in your repositories, the procedures are documented, your operator has already run them unaided, and the exit path was tested rather than promised.
We are deliberately built so that leaving is cheap. A relationship that survives because exit is expensive is not one we want.
Money, access grants, legal commitments, regulated decisions, and destructive operations are never delegated by default. An AI system cannot approve its own release, grant itself production access, or bypass external authorization for a consequential action.
In a first engagement we also avoid anything touching customers, money, or regulated decisions entirely — the first workflow is engineering-facing with an internal blast radius, because that is where you learn cheaply.
Yes, but expect us to say so. If you have no CI/CD, no infrastructure as code, and no cloud governance, there is nothing to install an agent runtime onto, and selling you one would be malpractice with an invoice attached.
That platform work is real work we know how to do, and we will scope it as its own engagement with its own shape. What we will not do is disguise it as an AI programme because that is what has budget this year.
We have published the failure modes, the graded evidence, the price bands, the exclusions, and the limits of our own capability. The first conversation is 45 minutes and one question: where are you on the ladder, and what has already gone wrong?