Enterprise AI agents are capable of doing impressive work.
Getting those agents to keep doing that work correctly once they enter a messy production environment, however, is proving to be a much harder task.
At VB Transform 2026, Bryan Silverthorn, Director of AGI Autonomy at Amazon, pointed to Cisco data showing that 85% of enterprises are piloting AI agents.
Of those enterprises, only 5% have shipped them to production.
That gap matters because an agent doesn’t necessarily have to crash to fail.
Silverthorn described a customer whose agent worked flawlessly for two months while handling software quality assurance tasks involving serial number extraction from screens.
It later began intermittently reading the wrong numbers after an underlying software change altered how the vision system behaved, even though the change was not obvious to humans.
Given this, Silverthorn broke reliability into four dimensions:
- Consistency
- Robustness
- Predictability
- Safety
This framework was designed to separate problems that often collapse into a single evaluation score, particularly because agents can perform consistently under controlled conditions while remaining fragile when their environment changes.
Krazimo, an AI development company co-founded by CEO Akhil Verghese, approaches the problem from a similar direction.
Its preference is to use deterministic workflows where possible, introduce constrained intelligence where judgment is required, and reserve fully autonomous agents for situations where the consequences of failure can be managed.
For Verghese, production readiness has less to do with whether an agent passed its initial tests and more to do with whether the organization knows what happens after those tests are over.
“We never consider a solution production-ready just because it passes its success criteria,” he says.
“Passing the criteria earns you the right to deploy; logging, monitoring, and regular checks and maintenance are what keep you deployed."
"No AI or machine learning system should ever be treated as launch-and-forget.”
In other words, organizations need to determine how reliably it can do the work, what happens when conditions change, and who is responsible when something goes wrong.
Design For Possible Failure Before Deployment
The most difficult AI failures are often the ones that look like normal operation.
In most cases, an agent that continues responding with incorrect information can remain unnoticed for weeks, particularly when the output looks reasonable enough to pass a casual review.
This is where Silverthorn’s and Krazimo’s approaches come in.
- Consistency asks whether the same input produces the same output.
- Robustness considers whether the system continues to perform when inputs change, such as different phrasing or unexpected formats.
- Predictability asks whether the organization can anticipate where the system will fail.
- Safety concerns itself with what happens when it does fail and whether the consequences remain contained.
What many organizations fail to understand, though, is that these are all different problems that respond differently to testing over time.
Pre-launch testing has limitations since teams can only test the variations they think to include.
But production systems will inevitably encounter changes that were unaccounted for during pre-launch testing.
Krazimo has encountered this problem with agents that answer questions using continuously updated knowledge graphs.
The agent itself continued operating correctly. The problem was that the source it was using had accumulated contradictory information over time.
That creates a particularly difficult failure mode.
An agent can faithfully retrieve information from its source and still give the wrong answer because the underlying source has become unreliable.
To solve this, Krazimo built monitoring that checks new information against existing data for contradictions and alerts the named human owner when conflicts appear.
So when a policy genuinely changes, the correction can be propagated through the knowledge graph rather than leaving conflicting versions in place.
“The important part is that none of that was in the original success criteria. It came from accepting that the system would be wrong later even though it was right at launch,” Verghese adds.
Build Reliability Into The Architecture
When AI behavior cannot be perfectly predicted, the system surrounding the model has to control the consequences.
Given this, Verghese advises teams to build reliability into the architecture from day one.
To achieve this, he recommends teams do the following:
1. Monitor outcomes instead of chasing perfect prediction
Krazimo focuses on a small set of north-star metrics that remain meaningful even when the underlying inputs vary.
If a conversion rate normally sits around 25% and falls below 20% for three consecutive days, the organization doesn’t need to know exactly which unseen condition caused the decline before taking action.
Instead, the system should flag the change and trigger an investigation.
That distinction matters because some AI failure modes are simply too difficult to enumerate in advance. Teams can spend months trying to anticipate every possible way an agent might behave differently.
As such, the more practical approach is to identify the outcomes that matter most and continuously watch them.
Start by asking these questions:
- Is the agent producing the expected business outcome?
- Has accuracy moved outside its normal range?
- Has the frequency of exceptions changed?
- Are users behaving differently after interacting with the system?
- Has a critical dependency or data source changed?
- Has performance remained within the range required by the business process?
2. Give every failure a human owner
Every important metric should have a person responsible for reviewing significant changes, investigating the cause, and deciding what happens next.
That accountability becomes particularly important as organizations give agents more autonomy.
“An agent is like a junior-employee-in-training. Autonomy is earned through measured reliability, not granted at deployment. The difference is that you can’t let past success create false confidence,” Verghese says.
“A human's track record gives you some basis for expecting future performance, but AI behavior can remain jagged."
"The monitoring and checking have to stay in place even after a system has been performing well for months.”
3. Constrain autonomy to the risk
Not every task needs an autonomous agent.
Krazimo's approach is straightforward:
- Use deterministic systems when a deterministic system can do the job.
- Introduce specific, modular points of non-determinism when judgment is actually needed.
- Use autonomous agents when the process genuinely benefits from that level of flexibility.
For example, Verghese typically allows cold outreach to run autonomously.
“Agents research prospects, run the sequences, and book calls on my calendar without me approving each message. If one of those messages is slightly off, the cost is one stranger who was unlikely to reply,” he says.
However, he takes a more measured and involved approach in other use cases, such as during the development of JSTFYD, an AI-powered insurance claim management system.
The system’s workflow produces a computation that justifies an expense in response to an insurance company's question.
So an incorrect answer could affect whether a client receives money they are legally owed or create a much more serious compliance problem.
That’s why JSTFYD has human approvals built into the workflow. The agent provides the answer along with its reasoning and evidence, and a person makes the final decision.
Earn The Right To Ship An Agent
The path out of enterprise AI's pilot phase begins before production engineering.
Krazimo's first filter is deciding whether a project is worth pursuing based on the reliability the specific process requires.
That may sound obvious, but vague definitions of success can allow a technically interesting project to continue long after the required level of performance is determined to be unrealistic.
According to Verghese, the best way to avoid this trap is to ask three simple questions:
- What does success actually mean for this workflow? The team needs a measurable definition tied to the business outcome, rather than a general statement that the agent should be useful or accurate.
- What level of performance would make the system useful? A process that requires 99.99% reliability shouldn't be evaluated using the same standard as an internal workflow where occasional errors can be reviewed and corrected.
- Can today's technology realistically reach that threshold? Small proofs of concept and prototypes can test the assumption before the organization commits to a full production build.
“We use our experience to decide which projects to take on. Often, the biggest savings often come from skipping ones where your intuition says today’s AI won’t meet the reliability the process demands,” Verghese adds.
Make Reliability The Price Of Admission
Enterprise AI has spent plenty of time proving that agents can do remarkable things. That made sense because AI was trying to carve out a place in the market at the time.
The next phase is less glamorous, however, is less glamorous, because businesses will have to decide which of those remarkable things are safe to trust.
That may mean fewer agents make it into production than the current AI hype cycle would suggest.
On the other hand, it may also mean the ones that do earn their place will have something more valuable than impressive demos behind them.







