Why AI in Document-Heavy Workflows Fails Without the Right Foundation - with Sumedh Chaudhary of IBM artwork

Why AI in Document-Heavy Workflows Fails Without the Right Foundation - with Sumedh Chaudhary of IBM

The AI in Business Podcast

June 24, 2026

Enterprise AI initiatives consistently break down in document-heavy environments, not because the underlying models are inadequate, but because fragmented data silos, page-break context loss, and uncoordinated extraction tools erode the semantic layer AI needs to reason accurately.
Speakers: Daniel Faggella, Sumedh Chaudhary
**Daniel Faggella** (0:12)
Welcome everyone to the Emerge AI in Business Podcast. Today's guest is Sumedh Chaudhary, CTO of US Industry Market at IBM. In this episode, Sumedh explains why so many enterprise AI deployments struggle with document-heavy workflows. Not because the technology falls short, but because critical context gets lost between pages, across data silos, and between disconnected tools. He outlines how a multi-agent architecture gives enterprises a reliable foundation they need to make document AI work. Today's episode is sponsored by Arango. A quick note for our audience that the views and opinions expressed by Sumedh Chaudhary on today's program are his own and do not reflect those of IBM or its leadership. In this episode, we cover how enterprises can build multi-agent AI architecture to handle document-heavy workflows, and the governance frameworks that determines whether those deployments scale. To go deeper on this topic and to learn how to structure landing pages for higher conversion, and how to use self-qualification systems to prioritize high-intent leads, download our free PDF report, B2B AI Lead Generation Guide at emerj.com/aig1.
That's emerj.com/aig1 to download your copy. Now the conversation with Sumedh.
Welcome to the show, Sumedh.

**Sumedh Chaudhary** (1:52)
Good morning. Thank you so much for having me on your show.

**Daniel Faggella** (1:55)
Absolutely. It's a pleasure to have you here. I'm going to dive right into it. There's no shortage of AI ambition in large enterprises at the moment. But the ROI conversation is definitely getting harder. Before we start getting into solutions for that, I'd love to start with the problem itself. We see in large enterprises with document-heavy or highly-regulated workflows, where do you usually see the biggest breakdown when organizations try to apply AI and how does it work at scale?

**Sumedh Chaudhary** (2:23)
Yeah, absolutely. I think this is a very common problem, and I think it's getting separated because of some of the changes in the technology landscape. As we all know, in the world of structured data, ability to do analytics, the ability to apply AI was the first victory for the technology industry.
As we evolved and moved into the unstructured data, we had a lot of success in the recent times using some of the generative AI models to be able to analyze and put solutions for unstructured data. But a document-heavy workflow presents the third level of challenge. This is a place where not only are you dealing with unstructured data, you're looking into a document file, which is quite complex. You're not looking at just unstructured data. You've got images in it, you've got tabular data in it, and there are something we call as page breaks. When you flip from one page to the next page, the computer and the systems that are designed to understand unstructured data somehow break down because they lose the semantic layer between the context of the data going from page 1 to page 2 That's one of the complex challenges of working with the document-heavy workflow.

**Daniel Faggella** (3:46)
Am I hearing you correctly that a lot of these failures come from the fact that all these different documents, formats, and different types of data all live in separate systems. If they were actually stored in one connected place, AI would have the context it needs to operate in a more reliable environment.

**Sumedh Chaudhary** (4:04)
Absolutely.
If I can give an analogy, if you've got a chef who is preparing a meal, if that person is trying to look at the ingredients that they have at hand, let's say they have spices, and let's say they have tools, if none of the spice jars are labeled, if you don't know how to recognize between the different spices, it could be very complex for even an experienced chef to prepare that absolutely beautiful meal. But that's one of the challenges is if you've got these pockets and silos of data living in your organization and enterprise that is spread over different silos in your organization, it becomes quite complex for even a very well-designed AI system to be able to put it all together.

**Daniel Faggella** (4:53)
That was a very practical analogy and very easy to understand and identify. Getting into the identification, we might have leaders in enterprises that aren't as involved in the AI functions. How would they be able to identify this from the outside? What does the leader actually see or feel when this is happening in their organization that there's no context?

**Sumedh Chaudhary** (5:13)
Yeah. I honestly work with these leaders every day. One of the things I've noticed over the past six to 12 months is, there is a little bit of a race to get to the finish line. I think everybody's trying to establish early success. Folks are looking at the low-hanging group. There are a lot of vendors in this space who are trying to target these enterprise leaders to highlight their products and solutions.

18 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000774043454