Track 2 · AI production engineering and AgentOps

Move one valuable AI-agent workflow from prototype to controlled production.

Peak Consulting is testing an Azure-based productionization service for companies whose agent prototypes are blocked by security, data access, evaluation, deployment, cost, observability, or unclear operating ownership.

1 named workflow 10–15 days readiness baseline 8–10 weeks launch project Human approval for consequential actions

Scope and prices are working hypotheses requiring separate buyer evidence.

The demo works. The production system does not exist yet.

Agentic systems are not only models and prompts. They combine business workflow, data, tools, authorization, application behavior, evaluation, telemetry, cost, and human intervention.

Control gap

Actions outrun permissions

The prototype can call tools, but nobody has defined what it may read, change, approve, retry, or never do.

Evidence gap

Demos replace evaluation

New prompts, tools, models, and indexes are judged by a few conversations rather than representative tests and release thresholds.

Operating gap

No owner can explain failure

Tracing, privacy, cost, escalation, rollback, and disablement are unresolved when the workflow reaches real users.

Prototype blocked by security Untracked agent tools No evaluation gate Uncontrolled trace data Cost and latency surprise

Bound the workflow before building the platform

The opening project productionizes one funded use case with named owners and reversible or human-approved actions. Additional agents and tool domains are separate decisions.

Step 1 · 10–15 business days

Agent Production Readiness Baseline

$15k–$25k working range

Assess the workflow, Foundry architecture, models, data, tools, identity, evaluation, tracing, privacy, cost, human approvals, and incident controls.

Output: target architecture, risk register, evaluation plan, control boundaries, and roadmap.

Step 2 · 8–10 weeks

Agent Launch Platform Project

$35k–$65k one workflow

Implement client-owned Azure/Foundry environments, deployment, approved knowledge and tools, evaluations, tracing, cost controls, approval, rollback, and runbooks.

Acceptance: the bounded workflow passes release gates and operates with a tested kill switch.

Step 3 · Fixed bridge

90-day Agent Stabilization

$8k–$15k/mo working range

Review traces and feedback, improve behavior through versioned changes, manage cost/latency, refine approval thresholds, and transfer ownership.

Not unlimited agent development or autonomous ownership of the business process.

The agent is only one layer of the production system

Microsoft Foundry supplies models, agents, evaluations, and governance capabilities. The workload still needs business boundaries, application controls, Azure foundations, and an operating system.

Business workflow
Owner Users Outcome Allowed actions Prohibited actions
Agent application
Orchestration State Instructions Approval experience
AI capabilities
Foundry models Agent Service Knowledge Approved tools
Azure controls
Projects and RBAC Managed identity Private network Policy and secrets
AgentOps
CI/CD Evaluations Traces Metrics and cost Incidents Rollback

The agent may propose an action. An external authorization control decides whether that identity may execute it, and a human approves consequential actions.

— Initial autonomy boundary

Ten questions must have evidence before launch

01

Outcome

Who owns the result, who uses it, and what improves?

02

Scope

What may the agent do, never do, and escalate?

03

Data

Which sources, identities, sensitivity, freshness, and retention apply?

04

Tools

Are permissions minimal, inputs validated, actions auditable, and retries safe?

05

Quality

Which representative cases and thresholds block a weak release?

06

Safety

How are injection, abuse, unsafe outputs, and prohibited actions handled?

07

Platform

Are environments, IaC, network, secrets, policy, and capacity repeatable?

08

Operations

Can owners trace, alert, intervene, roll back, and disable the workflow?

09

Economics

What are latency and model/tool cost per completed task?

10

Launch

Who signs off, supports the release, and reviews measured value?

Build the evidence and controls in the same delivery loop

PhaseHuman decisionEngineering outputAcceptance evidence
Week 0–1Workflow, owners, prohibited actions, and success casesCharter, risk register, evaluation dataset v1Business, technical, and security owners approve boundaries
Week 2Foundry, project, identity, and network decisionsIaC foundation and isolated non-production environmentEnvironment deploys repeatably
Week 3Model, grounding, and data accessModel and knowledge integrationSource access, freshness, and grounding tested
Week 4Tool and action boundariesLeast-privilege tool integrationsUnauthorized and malformed actions fail safely
Week 5–6Quality, safety, trace, and privacy thresholdsAutomated evaluations, tracing, and dashboardsWeak release is blocked; trace access/retention reviewed
Week 7Intervention and failure controlsHuman approval, rollback, disablement, and runbooksConsequential action requires approval; kill switch works
Week 8–10Bounded release and value reviewProduction launch, tuning, cost controls, handoffOwner can operate, inspect, and stop the workflow

Start where actions are reviewable and reversible

Better starting points

  • Internal knowledge with source links
  • Support research and response drafting
  • Document intake and classification
  • Engineering and incident investigation assistance
  • Compliance-evidence collection
  • Recommendations requiring human confirmation

Avoid at launch

  • Autonomous production infrastructure changes
  • Unrestricted shell, database, or cloud access
  • Final financial, legal, employment, or healthcare decisions
  • Payments, deletion, or commitments without approval
  • Broad cross-system write identities
  • Open-ended multi-agent autonomy without clear ownership

AI may help build the system. It cannot approve itself.

AI can accelerate
  • Architecture and IaC scaffolds
  • Evaluation-case and edge-case drafts
  • Trace, cost, and failure-pattern analysis
  • Prompt/tool documentation
  • Runbook and threat-scenario drafts
Humans must own
  • Business acceptance criteria
  • Trust boundaries and allowed actions
  • Production identities and permissions
  • Privacy, retention, and regulatory decisions
  • Release approval and incident response
Tracing is customer data.

Agent traces may contain prompts, responses, retrieved context, tool calls, intermediate steps, personal information, latency, tokens, and errors. Enable tracing deliberately; redact sensitive data; restrict access; and define retention before production.

Current references: Microsoft Foundry architecture, Foundry landing-zone baseline, tracing and data handling, and Agent Service transparency guidance.

AgentOps begins with the same Azure foundation as platform engineering.

Identity, network, IaC, CI/CD, observability, cost, recovery, and client ownership remain the common operating core.