

You are not hiring an enterprise AI agent deployment consultant to build an agent.
The agent itself is the smallest, cheapest and most replaceable artefact in the whole engagement. What you are buying is a set of decisions about scope, a set of boundaries about ownership, and a functioning operating model that survives after the consultant's last invoice clears. If the engagement produces a working agent and no operating model, you have bought a demo with better uptime.
The data supports that framing more directly than most consultants would like. Forrester's root-cause analysis of agent deployments reporting negative ROI at twelve months attributes 41% of those failures to unclear success criteria, 33% to insufficient tool or data access, and 26% to drift in evaluation coverage. Read the list again. Not one of them is a model-quality problem. They are scoping and ownership problems, which is to say they are problems that were somebody's job to solve before anyone wrote code.
This piece sets out what should actually sit on the consultant's side of that line, what should sit on yours, and where the seams between them tend to fail.
Six things, and the useful version of this list is written as a boundary rather than a capability statement. A capability statement tells you what a firm can do. A boundary tells you who gets called when something breaks.
Division of Responsibilities in Enterprise AI Agent Deployment.
The row that matters most is the last one. ISO/IEC 42001, the international AI management system standard, requires a documented owner responsible for each AI system's behaviour and outcomes. That requirement exists precisely because the failure it prevents is so common: a system in production that nobody is accountable for, discovered at the moment it does something expensive.
Worth stating plainly, because vague scope cuts both ways. A consultant should not own your data quality, your identity model, your change management, or the business decision about which process to automate. Firms that accept those responsibilities are either charging for a much larger engagement than you think, or they are about to miss them.
Not in the middle of the workstream. The deployment challenges that actually derail engagements sit on the seams between owners, where a task is adjacent to two people's jobs and therefore nobody's.
An agent acting on behalf of a user needs an identity, permissions and an audit trail. Whose identity does it act under? A service account with broad rights is easy and creates an audit problem.
Per-user delegation is correct and requires work in your identity provider that nobody scoped. This surfaces around week six in most engagements, and it is the single most common cause of slipped timelines after data readiness.
The most frequently reported source of delay in deploying AI agents in the enterprise is looping back to fix data readiness after skipping it.
It is skipped because it belongs to a team that is not in the deployment meeting, and it is discovered because the agent cannot answer questions the demo answered easily. The demo ran on curated documents. Production runs on what you actually have.
Who gets paged when the agent starts failing at 2am? The consultant is gone. The platform team did not build it. The business team cannot debug it. Deciding this before go-live takes twenty minutes. Deciding it during an incident takes a quarter of goodwill.
Models deprecate. Prompts drift as your product and policies change. Somebody owns re-running the evaluation suite when the underlying model version moves, and if that person is not named in the contract, the answer is nobody, and the agent quietly degrades over the following two quarters.
Agent spend is usage-based and grows with success. The median enterprise LLM bill has been reported growing more than seven times year on year entering 2026.
Whoever owns that budget line needs alerting thresholds and a per-run cap, decided during architecture rather than after the first surprising invoice.
Honest ranges, which vary far more by integration count and regulatory exposure than by agent complexity.
Realistic Deployment Timelines for Enterprise AI Agents.
Two figures worth setting against those. Median time-to-value across enterprise agent deployments is reported at roughly 5.1 months by BCG and Forrester, with function-level variation from about 3.4 months for sales development agents to 8.9 months for finance and operations.
And 88% of pilots never reach production at all, a figure originating in Anaconda and Forrester research and since replicated independently by a16z and the MIT Sloan CIO panel.
The gap between those two numbers is the whole subject of this article. Plenty of pilots work. Far fewer become systems someone runs.
A six-week quote is achievable if security review, identity work, data remediation and governance evidence are all somebody else's problem, or assumed to be zero. It usually means the proposal has scoped the build and excluded the deployment, which is the part the word deployment refers to.
Ask any firm quoting a short timeline which of the four they have included. The answer tells you what you are buying, and it is the fastest way to tell a deployment proposal from a build proposal.
Three matter, they overlap heavily, and none of them was designed for agents.
ISO/IEC 42001:2023 is the international AI management system standard, and the only one of the three offering certification, through a two-stage audit with an accredited body. It shares ISO's harmonised structure with ISO 27001 and ISO 9001, so it integrates with management systems you probably already run. Enterprise procurement teams increasingly require it from AI vendors as a purchase condition, which makes it a commercial asset rather than only a compliance one.
NIST AI RMF 1.0 is voluntary, has no certification path, and organises risk work into four functions: govern, map, measure and manage. Its influence exceeds its formal status, since US federal agencies reference its principles and enterprise buyers use it as a maturity benchmark.
The EU AI Act is binding regulation with a risk-tier classification. High-risk obligations were extended to 2 December 2027, which sounds distant and is not, given conformity work runs in quarters.
The common pattern in practice: ISO 42001 for the management system because customers ask for it, NIST AI RMF as the internal operating structure, and EU AI Act obligations layered on where jurisdiction and risk tier require them.
Because the frameworks overlap substantially, one control set can satisfy all three if it is designed that way from the start rather than retrofitted three times.
None of them was written for autonomous agents. NIST AI RMF predates production-grade agents. ISO 42001 is a management system standard rather than a technical one. The EU AI Act classifies systems rather than behaviours.
All three ask for something an agent system must be deliberately built to produce: evidence that oversight actually operated, not merely that a policy existed. Logs showing what happened are not the same as records showing which rule fired, on which call, with what outcome. That distinction is an architectural decision, and it is far cheaper to design in than to reconstruct.
For agentic AI enterprise deployment, the practical control set that satisfies most of what the three frameworks ask for comes down to five domains: policy articulation, access controls, observability, incident response, and drift monitoring. Build those and the framework mapping becomes documentation rather than engineering.
The handoff is the deliverable. Everything before it is work in progress.
A complete one produces six artefacts, and if any are missing the engagement is not finished regardless of whether the agent works.
The runbook. What the agent does, what it must never do, what its failure modes look like, and what to do about each. Written for the person who inherits it, not for the person who built it.
The evaluation suite, with ownership. The regression test set, the rubric, the pass thresholds, and a named person responsible for re-running it after any model or prompt change. This is the core of lifecycle management, and it is the artefact most often skipped.
The escalation path. Who is paged, at what threshold, with what runbook entry. Tested before go-live rather than during the first incident.
The governance evidence pack. The impact assessment, the access model, the decision records, mapped to whichever framework applies. Assembled during the build, not reconstructed afterwards.
The cost model and its alerts. Expected spend per run, the thresholds, the caps, and who owns the budget line.
Trained operators. Not a demo session. The people who will run it, having run it, under supervision, against real cases.
The mechanism that makes a handoff hold: the consultant stays available but stops driving, for around thirty days after go-live. The client's team runs the system, handles the first incidents, and calls when stuck. Any gap in the artefacts above surfaces in that window while someone who understands the system is still reachable.
Engagements that end at go-live look cheaper and are not.
Not the model, and not usually the build.
The pilot underestimates the programme. Moving from pilot to production commonly requires 250% to 400% more investment than the pilot itself, driven by data pipeline work, security hardening and integration depth. A pilot budget is not a fifth of a production budget; it is closer to a quarter of the next phase alone.
The build is a minority of lifetime cost. Initial build typically accounts for 25% to 35% of three-year total cost of ownership. Integration per system, inference, governance tooling and change management make up the rest, which is why AI agent deployment tools for enterprises should be evaluated on what they cost to operate rather than to adopt.
Deployment mode changes the shape. Private deployment and bring-your-own-key (BYOK) arrangements raise infrastructure and operational cost while satisfying residency and data-control requirements that would otherwise block the project. Kubernetes-based deployment adds platform overhead and buys portability and scaling control. These are not optimisations, they are constraints, and they should be settled in architecture rather than negotiated late.
Inference is usage-based and grows with success. The more the agent works, the more it costs. Budget for the successful case, not the pilot case.
The useful question for any proposal is not what it costs to build. It is what it costs to run in year two, and which of those lines the proposal has included.
Here is the part that cuts against hiring anyone, including us.
A consultant is worth engaging when the constraint is capability you do not have and will not build: agent architecture, evaluation design, governance evidence, deployment patterns. A consultant is not worth engaging when the constraint is organisational. If your data is not accessible, your identity model is unresolved, or nobody senior has decided which process to automate, an external team will discover those problems at your expense and hand you a report you could have written internally.
The uncomfortable version: a meaningful share of what makes enterprise AI agent deployment succeed is not deployable by an outside party at all. Naming an owner, granting access, agreeing a success metric and committing to run the thing are internal decisions. A good consultant will force them early. A good consultant cannot make them for you.
If you are not willing to name the owner before go-live, the honest recommendation is to wait until you are, because the alternative is a working system that decays quietly and a post-mortem about the vendor.
Codiste deploys agentic systems into enterprise environments and hands them over with the runbooks, evaluation suites and governance evidence that make them survivable. If your last pilot worked and never reached production, our AI agent development services start by finding out which seam it fell through.




Every great partnership begins with a conversation. Whether you're exploring possibilities or ready to scale, our team of specialists will help you navigate the journey.