Article

AI Agent Governance: Who Answers When an Agent Gets It Wrong?

Summarize with

By Jess Lampe, Global Lead Technologist, Launch Consulting Group

Most customer-facing AI rollouts have a plan for what the agent will do. Very few have a name attached to what happens when it does the wrong thing. This is the central challenge of AI agent governance.

We see this gap constantly in the deployments that cross our desks: not the model, not the interface, not the deployment timeline, but the actual exposure sitting inside most AI roadmaps right now.

Organizations are shipping autonomous decision-making into conversations with customers and treating accountability as something that will sort itself out later, the way it never quite did with any other system that made decisions without a person in the loop.

Our position at Launch is blunt, and it's not a new one: governing an AI agent is not a new discipline. It's the same one organizations have always needed for people and automated systems that don't behave perfectly every time, now applied to something that talks directly to customers. “The AI did it” was never going to be a valid answer. The agent acts. A person still has to answer for it. That principle is the foundation of AI agent accountability.

The cost that never made it into the build budget

Most organizations cost the build and stop there. In every engagement where something's gone sideways, the real cost showed up later, in the run, in places a project dashboard doesn't reach.

Customer trust erodes quietly and takes far longer to rebuild than it took to lose. Every wrong answer generates rework somewhere downstream causing verification labor that's real cost, not overhead. A wrong answer that touches a regulated decision turns into liability exposure, not just a bad interaction.

And the version we see catching people off guard most often is the one nobody sees coming: a system that was accurate on launch day, degrading quietly as the model and the world around it shift, unnoticed because nobody instrumented for it.

None of this is a technology bill. It's an operating bill. It doesn't show up until someone's already paying it.

What separates the organizations that don't get burned: AI agent governance in practice

Across the deployments we've been closest to, four disciplines show up, in this order, in every customer-facing AI rollout that hasn't had an incident yet.

1. Name the owner before you name the model

“The AI did it” is not a valid answer.

The agent acts. A person answers for it. Every customer-facing agent needs a single named owner accountable for its behavior, its value, and its lifecycle the same way a product manager owns a product, not the way a committee owns a policy. That owner instruments the monitoring; they don't just approve the launch and move on.

We've watched what happens when that step gets skipped: deploying an agent without naming that owner isn't efficiency, it's a transfer of decision rights with nobody left to answer for the decisions. Practically, that means a documented RACI behind every customer-facing system, an escalation path that reaches a human fast, and a contestability process so a customer who got a wrong answer can actually challenge it.

Why this comes first: every other control in this list is something the owner is supposed to be running. Skip this step and the rest become nobody's job.

2. Tier the controls to what the agent can actually do

A read-only summarizer and a system that touches a customer's money are not the same risk. Stop governing them identically.

The most common governance mistake we run into isn't too little control, it's the same control everywhere. Match guardrails, human-in-the-loop review, and testing intensity to what the agent can actually do to a customer, not to how impressive the use case sounds in a steering committee deck. This is where human oversight in AI becomes operational rather than theoretical.

That includes rethinking what “testing” means. Customer-facing AI needs behavioral validation, not traditional QA: hallucination and abuse testing, reasoning regression, champion/challenger comparisons, human review gates on the highest-stakes paths. Bias and accuracy aren't things you certify once at launch. They're behaviors you test for continuously, because the model that passed testing in March isn't guaranteed to behave the same way in September.

The other half of tiering is identity, not just review. Each agent should carry its own identity and the minimum access it needs to do its one job then audited like any other actor with system access, not granted broad access because it was faster to set up that way.

3. Baseline before you build

No baseline, no benefit.

This is the mistake we see undo more otherwise-solid AI investments than any other: if you didn't measure the process before the AI touched it, you can't prove the AI improved it — and you won't be able to tell when it quietly stops improving it. There's no “before” left to compare the “after” against, and every later argument about ROI or trust turns into a story instead of a number.

Once the baseline exists, measure a small set of signals per use case, split into two categories that both matter: operational quality (accuracy, resolution rate, cost per transaction, throughput) and customer experience (satisfaction, escalation and handoff-to-human rates, repeat-contact rates). Treat every human correction as signal, not cleanup knowing that a rising correction rate is an early warning that trust is slipping, and it shows up before CSAT does.

Fund customer-facing AI the way you'd fund anything else you can't fully predict yet: in short, evidence-based increments tied to a measured outcome, not a multi-year commitment made on a demo.

4. Watch it after launch, not just at launch

A system that looks great in week one can still be decaying by week twelve.

This is the discipline we see skipped most often, because governance gets treated as a gate to clear before launch, rather than a thread that travels with the system for its whole life. We've watched what happens when it's bolted on after the fact instead: agents drift, cost creeps, quality degrades unwatched, and one incident freezes a program that nobody can vouch for anymore. This is why AI agent risk management has to continue after deployment, not end when the agent clears its launch review.

Instrumenting for behavioral drift, safety signals, and reasoning traces from day one is what turns a good launch into a durable one. A strong first week proves nothing on its own. Only sustained performance against the baseline does.

The floor is rising, and it's expensive

What we're watching regulators converge on is a consistent direction: regulate the use, not the model. Obligations increasingly attach to what an AI actually does in a given context — especially consequential decisions that affect consumers — which is exactly why tiering controls to risk is becoming table stakes rather than best practice.

The stakes are no longer theoretical. The EU AI Act carries fines up to €35 million or 7% of global revenue. In the US, state-level laws like the Colorado AI Act are attaching per-consumer penalties, in the range of $20,000 each, to deceptive automated decisions, which means US organizations can't afford to wait for a single federal standard to settle before acting.

Preparing for this doesn't require a new department. It requires four things most organizations don't have yet:

  • an AI inventory that includes shadow AI (you can't govern what you can't see, and most organizations already have employees running customer work through tools nobody approved)
  • a live agent registry with named owners and evidence
  • vendor and model standards written into contracts
  • an incident playbook paired with the contestability process from discipline one

Third-party model risk is arguably higher than traditional vendor risk, because the model behind the contract keeps changing after the ink dries.

Where to start with AI agent governance

The choice in front of most tech leaders isn't whether to deploy customer-facing AI faster or slower. It's whether ownership gets named now, while the AI already in production is still ungoverned, or gets sorted out later, after an incident forces the question.

Five places to start:

  1. Name an owner for every customer-facing agent already in production. If no one can be named, that's the first gap to close.
  1. Find the shadow AI. Assume customer-facing tools are already in use that nobody approved and go find them before a regulator does.
  1. Baseline the process the AI is touching, in the metrics that matter to the business, before extending the deployment further.
  1. Tier the controls on what's live today instead of applying one policy to every use case.
  1. Build the contestability path for the way a customer reaches a human when the agent gets it wrong, and build it before it's needed, not after.

An agent without a named owner isn't a shortcut. It's a decision nobody's actually responsible for. At Launch, we’d rather see that fixed before the next rollout than after the first incident.

With a career spanning applied AI, global consulting, cloud architecture, and platform delivery, Jess Lampe leads Launch’s Technology Council and serves as the firm’s Global Lead Technologist. In his role, he focuses on aligning tech expertise across teams and supporting consistency in how digital and AI-driven transformations are delivered to Fortune 1000 clients.

Back to top
Table of Contents
Back to top

By Jess Lampe, Global Lead Technologist, Launch Consulting Group

Most customer-facing AI rollouts have a plan for what the agent will do. Very few have a name attached to what happens when it does the wrong thing. This is the central challenge of AI agent governance.

We see this gap constantly in the deployments that cross our desks: not the model, not the interface, not the deployment timeline, but the actual exposure sitting inside most AI roadmaps right now.

Organizations are shipping autonomous decision-making into conversations with customers and treating accountability as something that will sort itself out later, the way it never quite did with any other system that made decisions without a person in the loop.

Our position at Launch is blunt, and it's not a new one: governing an AI agent is not a new discipline. It's the same one organizations have always needed for people and automated systems that don't behave perfectly every time, now applied to something that talks directly to customers. “The AI did it” was never going to be a valid answer. The agent acts. A person still has to answer for it. That principle is the foundation of AI agent accountability.

The cost that never made it into the build budget

Most organizations cost the build and stop there. In every engagement where something's gone sideways, the real cost showed up later, in the run, in places a project dashboard doesn't reach.

Customer trust erodes quietly and takes far longer to rebuild than it took to lose. Every wrong answer generates rework somewhere downstream causing verification labor that's real cost, not overhead. A wrong answer that touches a regulated decision turns into liability exposure, not just a bad interaction.

And the version we see catching people off guard most often is the one nobody sees coming: a system that was accurate on launch day, degrading quietly as the model and the world around it shift, unnoticed because nobody instrumented for it.

None of this is a technology bill. It's an operating bill. It doesn't show up until someone's already paying it.

What separates the organizations that don't get burned: AI agent governance in practice

Across the deployments we've been closest to, four disciplines show up, in this order, in every customer-facing AI rollout that hasn't had an incident yet.

1. Name the owner before you name the model

“The AI did it” is not a valid answer.

The agent acts. A person answers for it. Every customer-facing agent needs a single named owner accountable for its behavior, its value, and its lifecycle the same way a product manager owns a product, not the way a committee owns a policy. That owner instruments the monitoring; they don't just approve the launch and move on.

We've watched what happens when that step gets skipped: deploying an agent without naming that owner isn't efficiency, it's a transfer of decision rights with nobody left to answer for the decisions. Practically, that means a documented RACI behind every customer-facing system, an escalation path that reaches a human fast, and a contestability process so a customer who got a wrong answer can actually challenge it.

Why this comes first: every other control in this list is something the owner is supposed to be running. Skip this step and the rest become nobody's job.

2. Tier the controls to what the agent can actually do

A read-only summarizer and a system that touches a customer's money are not the same risk. Stop governing them identically.

The most common governance mistake we run into isn't too little control, it's the same control everywhere. Match guardrails, human-in-the-loop review, and testing intensity to what the agent can actually do to a customer, not to how impressive the use case sounds in a steering committee deck. This is where human oversight in AI becomes operational rather than theoretical.

That includes rethinking what “testing” means. Customer-facing AI needs behavioral validation, not traditional QA: hallucination and abuse testing, reasoning regression, champion/challenger comparisons, human review gates on the highest-stakes paths. Bias and accuracy aren't things you certify once at launch. They're behaviors you test for continuously, because the model that passed testing in March isn't guaranteed to behave the same way in September.

The other half of tiering is identity, not just review. Each agent should carry its own identity and the minimum access it needs to do its one job then audited like any other actor with system access, not granted broad access because it was faster to set up that way.

3. Baseline before you build

No baseline, no benefit.

This is the mistake we see undo more otherwise-solid AI investments than any other: if you didn't measure the process before the AI touched it, you can't prove the AI improved it — and you won't be able to tell when it quietly stops improving it. There's no “before” left to compare the “after” against, and every later argument about ROI or trust turns into a story instead of a number.

Once the baseline exists, measure a small set of signals per use case, split into two categories that both matter: operational quality (accuracy, resolution rate, cost per transaction, throughput) and customer experience (satisfaction, escalation and handoff-to-human rates, repeat-contact rates). Treat every human correction as signal, not cleanup knowing that a rising correction rate is an early warning that trust is slipping, and it shows up before CSAT does.

Fund customer-facing AI the way you'd fund anything else you can't fully predict yet: in short, evidence-based increments tied to a measured outcome, not a multi-year commitment made on a demo.

4. Watch it after launch, not just at launch

A system that looks great in week one can still be decaying by week twelve.

This is the discipline we see skipped most often, because governance gets treated as a gate to clear before launch, rather than a thread that travels with the system for its whole life. We've watched what happens when it's bolted on after the fact instead: agents drift, cost creeps, quality degrades unwatched, and one incident freezes a program that nobody can vouch for anymore. This is why AI agent risk management has to continue after deployment, not end when the agent clears its launch review.

Instrumenting for behavioral drift, safety signals, and reasoning traces from day one is what turns a good launch into a durable one. A strong first week proves nothing on its own. Only sustained performance against the baseline does.

The floor is rising, and it's expensive

What we're watching regulators converge on is a consistent direction: regulate the use, not the model. Obligations increasingly attach to what an AI actually does in a given context — especially consequential decisions that affect consumers — which is exactly why tiering controls to risk is becoming table stakes rather than best practice.

The stakes are no longer theoretical. The EU AI Act carries fines up to €35 million or 7% of global revenue. In the US, state-level laws like the Colorado AI Act are attaching per-consumer penalties, in the range of $20,000 each, to deceptive automated decisions, which means US organizations can't afford to wait for a single federal standard to settle before acting.

Preparing for this doesn't require a new department. It requires four things most organizations don't have yet:

  • an AI inventory that includes shadow AI (you can't govern what you can't see, and most organizations already have employees running customer work through tools nobody approved)
  • a live agent registry with named owners and evidence
  • vendor and model standards written into contracts
  • an incident playbook paired with the contestability process from discipline one

Third-party model risk is arguably higher than traditional vendor risk, because the model behind the contract keeps changing after the ink dries.

Where to start with AI agent governance

The choice in front of most tech leaders isn't whether to deploy customer-facing AI faster or slower. It's whether ownership gets named now, while the AI already in production is still ungoverned, or gets sorted out later, after an incident forces the question.

Five places to start:

  1. Name an owner for every customer-facing agent already in production. If no one can be named, that's the first gap to close.
  1. Find the shadow AI. Assume customer-facing tools are already in use that nobody approved and go find them before a regulator does.
  1. Baseline the process the AI is touching, in the metrics that matter to the business, before extending the deployment further.
  1. Tier the controls on what's live today instead of applying one policy to every use case.
  1. Build the contestability path for the way a customer reaches a human when the agent gets it wrong, and build it before it's needed, not after.

An agent without a named owner isn't a shortcut. It's a decision nobody's actually responsible for. At Launch, we’d rather see that fixed before the next rollout than after the first incident.

With a career spanning applied AI, global consulting, cloud architecture, and platform delivery, Jess Lampe leads Launch’s Technology Council and serves as the firm’s Global Lead Technologist. In his role, he focuses on aligning tech expertise across teams and supporting consistency in how digital and AI-driven transformations are delivered to Fortune 1000 clients.

Back to top
Launch Consulting Logo
Locations