Service catalogue

Four engagements, one alternative entry, and a module catalogue.

Every engagement is fixed-scope and fixed-fee against a bounded outcome. Every one names the rung it covers and the rungs it does not. Nothing here is sold by the hour, and nothing is sold as a retainer before there is something in production worth operating.

Twelve modular platform capabilities grouped across four planes and connected by a shared data spine.

Quoted as a sequence

Entry plus core is $105k–$190k. We say that in the first meeting. Discovering it in month three reads as a bait-and-switch, and it is.

Outcome, never hours

Fixed fee per bounded outcome. New priorities replace old ones or take a change order. Our efficiency is our margin problem, not your discount.

Every engagement is exit-ready

Artifacts live in your repositories and your CI. The acceptance test is your operator running the documented procedure without us in the room.

Agent Runtime & Control Baseline

Three to four weeks, fixed fee, working range $30k–$50k. Read-only first, so it clears security review without a fight. It is designed to be worth the money even if you never buy anything else — and roughly a third of its value is discovering that a thing you believed was true about your estate is not.

What we do

  • Estate inventory. Every AI and agent workload already running, its owner, its trigger, its blast radius, and its cost — including the ones nobody told you about
  • Tool plane and identity assessment. What agents can call, under whose credentials, with what scope, and what happens on a bad day
  • Lifecycle and containment audit. Including a live disable-survives-restart test on a non-production agent
  • Outcome ledger instrumentation. Flow, agent, adoption, and business tiers wired to real telemetry — before anything is built
  • Golden task suite. 20–40 representative engineering tasks from your team’s actual recent work, with known-good outcomes and deliberate false-completion traps
  • Data path review. Which datasets agents already read or write, classification gaps, and whether a masked test-data path exists at all

What you keep

  • A ranked gap report with a costed 90-day plan, prioritised by blast radius rather than by what we sell
  • The golden task suite, in your repository, runnable in your CI
  • The outcome ledger definitions and the baseline numbers
  • The agent registry as it stands, with named owners
  • Evidence of what the containment test actually did

When this is not right for you: if you have no platform — no CI/CD, no IaC, no cloud governance — there is nothing to install a runtime onto and the honest sequence is platform first. We will say so.

Runtime, Tool Plane & Guardrails

Six to eight weeks, working range $75k–$140k. This is the engagement a platform lead can describe to their own leadership in one sentence: agents can now do work here, inside limits that hold, and we can prove the limits hold.

Built

  • Environment fabric — ephemeral, isolated, fast environments agents can create, break, and destroy, with seeded and masked data. Provisioning time is a tracked SLO, because an agent waiting 25 minutes cannot iterate
  • Tool and action plane — the permissioned catalogue of what agents may do. Each tool a versioned contract with typed I/O, idempotency semantics, prohibited operations, a dry-run mode, and trace correlation
  • Identity and continuous authorization — per-agent, per-task scoped workload identity, short-lived credentials, operation-level authorization. No shared service principals, no standing production write access
  • Approval matrix — mandatory human approval for money, access, data egress, regulated decisions, and destructive actions
  • Agent registry and lifecycle control — every running agent, its owner, scope, supervisor, restart policy, and a disable path that survives redeploy and scheduler
  • Data contracts where data is in the path — because agents writing SQL against uncontracted tables is a control failure, not an analytics inconvenience

Acceptance evidence

  • A containment drill, run live, including a restart attempt — and it holds for a full cycle
  • Evidence that an unsafe action was correctly refused, not just that a safe one succeeded
  • Every write path produces an immutable audit record with trace correlation
  • Dry-run output for each destructive tool, reviewed before first live use
  • Environment provisioning meets its stated SLO from a cold start
  • Your operator runs the whole procedure once, unaided, and we watch
BB-2BB-4BB-5BB-8BB-9

The full technical detail →

Assurance & Context

Six to ten weeks, working range $80k–$150k. Guardrails filter inputs and outputs. No guardrail catches a plausible artifact sourced from the wrong place. This is the engagement that makes agent output trustworthy at volume, and it is where the relationship stops being replaceable by a product release.

Built

  • System context graph — a generated, continuously refreshed, machine-readable map of services, repos, pipelines, datasets, environments, dependencies, owners, SLOs, on-call, and cost centres. Built from IaC state, CI config, the cloud resource graph, catalogue metadata, git history, and incident records — never from a wiki somebody has to maintain
  • Execution memory — active plans with completion criteria, decision logs, progress files, quality grades, and structured handoffs, so a fresh session resumes instead of restarting
  • Provenance and postconditions — layer 0 of the evidence plane. Never accept a completion claim; assert independently that the end state holds, that inputs came from the authorised source at the expected version, and that prohibited side effects did not occur
  • Verification speed — test selection, parallelism, caching, flake quarantine, and merge-queue design, so time-to-verdict stays under the time an agent takes to produce the next change
  • Operator adoption track — the operators who own the workflow help author the golden suite, human judgment stays authoritative at named decision points, and override rates get acted on
  • One workflow in production — engineering-facing, internal blast radius, fast feedback

Measured, not asserted

  • Resumption success rate — the share of interrupted tasks a fresh session continues correctly without human re-briefing
  • False-completion catch rate — against traps we built into your golden suite
  • Time-to-verdict — baseline, target, and the flake budget that protects it
  • Cost per successfully completed task — the number a CFO can actually use
  • Override rate and week-four sustained usage — whether anyone still trusts it a month later
BB-1BB-3BB-6BB-7BB-12

The full technical detail →

Candidate first workflows. Dependency and framework upgrades at scale; test backfill and flake reduction; IaC drift remediation; data pipeline migration; incident-context assembly; SQL or dbt model generation behind contract gates. All engineering-facing, all with a buyer who feels the pain personally. Nothing touching customers, money, or regulated decisions in engagement one.

— How we pick the wedge workflow with you

AgentOps

$15k–$40k per month, offered only once you have an agentic workflow in production. We do not sell a retainer before there is something to operate — a subscription with nothing running under it is a tax, not a service.

Model and provider change

Evaluation of every model or provider change against your golden suite before it reaches your workflows. Models change under you whether or not anyone asked them to, and a silent capability regression is indistinguishable from a bug you caused.

Evaluation and drift

Eval maintenance, new traps as new failure modes appear, long-run reliability review, and data-quality gate operation on every agent-touched dataset.

Economics and incidents

Cost-per-successful-task management against a stated budget rule, agent incident review with a named human owner, and the adoption and override review that tells you whether anyone is actually using it.

When data is the whole engagement

Data is usually a component of the rungs above. Occasionally it is the problem, and starting anywhere else wastes your money. The standard door is still the runtime baseline — this is the exception, not a second front door.

The signs

  • An agent programme is blocked specifically on grounding quality or confidently wrong answers
  • There is no test-data path, so nothing can be safely built — very common, very rarely diagnosed
  • An incident traced back to stale or mis-permissioned data
  • Agents are already writing SQL or pipelines against an uncontracted warehouse

The engagement

A data readiness baseline establishes what agents can actually see and whether retrieval respects the permissions of the person asking. A foundation project then builds the semantic and contract layer, quality gates, lineage, and the masked environment path that unblocks everything else.

Working ranges: baseline $12k–$20k, foundation project $30k–$55k, 90-day stabilisation $6k–$12k/month. Sold alongside the ladder, not instead of it. See the data foundation →

Modules, components, and integrations

The catalogue is what stops a repeated need becoming one-off custom work. Anything marked destination is part of the published path and is built when a client pays for the behaviour — never speculatively, and never implied as present capability.

ModuleWhat it doesRungStatus
Outcome ledgerFour-tier measurement spine wired to real telemetry; cost per successfully completed task1Available
Golden task harnessClient-specific task corpus with known-good outcomes and false-completion traps; runs in your CI1Available
Agent registryEvery running agent, owner, scope, supervisor, restart policy, disable path1–2Available
Tool contract kitVersioned tool contracts, typed I/O, dry-run, idempotency, per-tool policy binding, audit2Available
Masked environment provisioningSeeded, masked, realistic data subsets for ephemeral agent environments2Available
Approval & action gatewayOne enforcement path for policy, approval, and consequential action2Available
Context graph builderGenerated agent-facing projection of the estate, ingested from what you already run3Available
Provenance & postcondition harnessLayer-0 evidence: independent end-state assertion, input provenance, side-effect checks3Available
Execution memory schemaPlans, decision logs, progress, structured handoffs, stale-rule hygiene job3Available
Semantic & contract layerGoverned metrics and dimensions as the agent’s data API, with CI enforcement3Available
Quality, lineage & observabilityFreshness and volume SLAs, OpenLineage wiring, anomaly detection, gates as merge gates3Available
Model portability boundaryProvider-abstraction boundary, routing policy as code, per-route evals, fallback and exit test3On request
Model gatewayFull routing gateway. Deliberately deferred until quality-adjusted cost improves after gateway fees and added operational complexity3+Deferred by design
Communication RouterPrivate-by-default agent communication with policy-driven promotion of material findings to humans — prevents collaboration overload4Destination
A2A interoperability gatewayCross-vendor agent-to-agent protocol translation and trust boundaries4Destination
Durable task & workflow serviceCanonical task record, durable state owned by the workflow engine rather than the model4Destination
Cross-functional workflow packOne end-to-end workflow spanning teams, vendors, and runtimes4Destination
Module catalogue & lifecycleRegistry, entitlement, release lifecycle, multi-business-unit governance5Destination

The 70/20/10 rule governs every module: 70% supported core, 20% configuration, at most 10% client-specific code. A permanent fork gets priced properly or declined — quietly maintaining bespoke forks is how good delivery organisations become bad ones.

The platform and data engineering we bring to the work

The agent-facing layer is new. What it stands on is not, and we do not treat it as an afterthought — most of what makes an agent programme survive contact with production is ordinary platform engineering done properly.

Cloud & platform

  • Landing zones, identity, network segmentation, policy
  • Infrastructure as code and drift control
  • Kubernetes, App Service, or Container Apps — chosen on workload, not fashion
  • Cost allocation and operating guardrails

Delivery & operations

  • CI/CD golden paths and release telemetry
  • Observability, SLOs, and actionable alerting
  • Incident learning and tested recovery paths
  • Rollback as a design requirement, not a hope

Data engineering

  • Ingestion and CDC, lakehouse layout, retention and time travel
  • Transformation and modelling with grain discipline
  • Semantic layer, contracts, and CI enforcement
  • Classification, masking, workload identity, cost attribution

Not sure which one you need? That is the first meeting.

Bring the agent nobody owns, the pipeline nobody can reproduce, or the audit finding you cannot answer. We will place you on the ladder and tell you what the next rung costs. If the honest answer is that you are not ready for any of this yet, we will say that instead.