Back

Watsonx and the AI Enterprise Reality Check: What Actually Happens When You Ship an AI Assistant to 500+ Internal Users

6 MINS

Watsonx and the AI Enterprise Reality Check: What Actually Happens When You Ship an AI Assistant to 500+ Internal Users

In 2023 and 2024, AI assistant demos were everywhere. Beautiful, fluent, confident. The bot answered questions perfectly. The stakeholders were impressed. The executive sponsor said "let's do this."

I was the PM who had to figure out what happened after "let's do this."

At IBM, I shipped a Watsonx.ai-powered chatbot to 500+ internal users across our CIO initiatives. I want to tell you what actually happened — not the deck version, the real version.

The confidence problem nobody talks about in the demo

In the demo, the AI gives a clean, accurate answer. In production, the AI gives a clean, confident answer that is sometimes wrong.

This is the central problem of internal AI assistants, and it's almost never discussed seriously in the planning phase. When a user asks the chatbot about the deal workflow and gets a slightly outdated answer delivered with the same tone as a correct answer, two things happen. First: they may act on it. Second: the next time they have a question, they're less likely to trust the bot — not because it was wrong, but because they can't tell when it's right.

Trust is asymmetric. It takes many correct answers to earn it and one confident wrong answer to break it. The moment a user says "I can't tell when to trust this thing," you've lost adoption and it won't come back without a significant redesign.

What we changed after the first wave of user feedback

We hadn't planned to revisit the confidence handling. But NPS was not moving the way we needed it to, and session feedback was consistent: users wanted to know when the bot was uncertain.

So we made the bot tell them. We built explicit uncertainty signaling — responses that acknowledged the limits of the available context, pointed to a human escalation path when the topic was high-stakes, and quoted the source of the information rather than synthesizing it into a confident-sounding statement.

Adoption went up. NPS moved. Not because the model got smarter — because users felt safer using it.

That lesson has stayed with me: AI product design is not just about capability, it's about calibration. The gap between what users expect from a bot and what it can reliably deliver is a product problem, not a model problem.

Change management is the actual work

The chatbot was one of four tools I was managing for IBM's CIO. The others — a Global Solutioning Tool, a Deal Workflow Tool, a Conga Salesforce integration — all taught me the same lesson in different ways: adoption is a product problem, not a training problem.

When we launched the Deal Workflow Tool, usage was low for the first six weeks. The tool worked. The training was solid. What we'd missed was that the people who used the old process didn't feel the pain we'd assumed they felt. They had workarounds. The workarounds were annoying but familiar. We had to make the tool tangibly better for their specific annoying scenarios — not just better in aggregate — before they'd commit to switching.

We ran targeted experiments. We tracked specific user segments. We found the two workflow moments where the new tool was dramatically faster and made sure those were the first things new users experienced. Adoption improved by 30% over the quarter.

That's product work. Not comms. Not training. Product.

What I'd tell any PM starting an enterprise AI initiative

Three things, plainly:

The first version will lack trust signals. Build them in before launch, not after the feedback comes in. Users forgive imperfection; they don't forgive being misled by something that sounds certain.

Adoption is a product metric, not a rollout metric. If people aren't using the tool after you've announced it and trained them, the product needs to change — not the comms.

The value has to be specific. "This saves time" is not enough. "This saves 12 minutes in the deal approval flow that everyone hates" — that's a reason to switch.

Enterprise AI is having its hype cycle. The PMs who navigate it well will be the ones who treat the adoption curve like any other product problem: measure it, hypothesize, test, and iterate. The model is the least interesting variable in the equation.

Background

Adarsh skipped presentations and built real AI products.

Adarsh Patnaik was part of the January 2026 cohort at Curious PM, alongside 13 other talented participants.