**SPEAKER_2** (0:06)
Are you trying to navigate the hype of AI? Now is the time to capitalize. Investment in AI has reached a new high with three quarters of CEOs already personally using AI tools. Gartner is here to help, providing executive leaders with expert insights, strategic guidance and actionable advice to make informed decisions and stay ahead of industry trends. Visit gartner.com or click the link in the show notes to download our AI Action Plan and learn how to build the right AI strategy for your organization.
**Karen Stokes-Lockhart** (0:47)
Welcome to Gartner ThinkCast. I'm Karen Stokes-Lockhart. This week, we're exploring one of the most persistent and evolving challenges with AI, hallucinations. If you've ever used any kind of AI tool like ChatGPT or Copilot, you've undoubtedly come across a few too many errors that are stated as fact. But why does this continue to happen? Featured as part of our Top of Mind Video Insights, Gartner Chief of Research Chris Howard will walk through a few uncanny examples, the techniques employed by neural networks to constrain and refine these answers, and even how hallucinations can actually be helpful. Now here's Chris.
**Chris Howard** (1:27)
Hi, everybody. Welcome to Top of Mind. I'm Chris Howard. When I started this series almost, gosh, two and a half years ago, now maybe even longer, it was started around the rise of interest in generative AI and has remained pretty consistent. And we've talked about AI a lot during Top of Mind. And it still stays really dominant. So I thought what I do is kind of talk a little bit about where we are, but some of the concerns that people still have and maybe demystifying some of those, especially around the topic of hallucination. So when generative AI really hit the public consciousness in the beginning of 2023, the end of 22, beginning of 23, people were fascinated by it, right? We talked about how people were camped out in the uncanny valley, kind of experiencing this interaction with this thing that seemed to be human.
And some of the funny bits about it were that it would make stuff up, right? So it's a prediction machine, essentially. So if it hit a spot where it really knew something should happen, but didn't have data to fill it, it would make something up. So for example, at Gartner, a number of analysts asked generative AI tools like ChatGPT to generate a biography for them. And it was very interesting what happened. So any public information around these analysts that existed had become part of the model would come out. So some of their background, their history, and it got most of it right with a few mistakes. But in almost every case, ChatGPT created an obituary for the analysts. So why? Why did that happen? Well, what we think is that most of the biographies it was trained on, so GPT was trained on, were historical biographies or maybe Wikipedia biographies, where most of them actually were from people who had passed. So it thought in its mind, in its predictive mind, there should be an obituary that's part of this biography. So it would create one. It would sort of make one up and say, unfortunately, Chris Howard died in 1993, blah, blah, blah, blah, blah. Funny, right? Not funny if what you're trying to do is to make a decision that informs a business strategy or makes a decision for a loan or something like that. So naturally, there was a lot of concern about generative AI and its hallucinations. Now, it has changed quite a bit since then. Because if you think about where hallucinations come from, it is a predictive machine filling in a space where it believes that something should go. It is drawing from its experience, from its training to do that. So part of reducing the hallucinations is to actually reduce the training space and to constrain it. Either by saying only use this data and maybe ignore this other type of data. So you're narrowing in on a set of data that it's going to draw from and it's going to draw its intelligence from. Or you could do things like using filters on the input or the output of the prompt to make sure that those citations it's using are real, and it's going back and interrogating the output of the prompt itself. Eventually what's happening, of course, is multi-agent systems bring that to another level where you have agents within a complex system that are actually working together to solve maybe a hard problem. For example, maybe it's a diagnosis problem in health care and you had multiple agents interpreting, say, the records from a patient to try to determine a diagnostic path. Now, in a real hospital setting, this is done by panels of experts. Some of you know that I'm a cancer survivor, so I survived stage 4 lymphoma about eight years ago. And what I learned, because I was interested while I was going through it, like, what am I experiencing? That there was a tumor board at Yale New Haven Hospital that met regularly to discuss complex cases. And they would all come from slightly different positions, some from a diagnostic position or a clinical position, some from testing, but there are others that represented, say, the financial side of like, how much are these tests going to cost and so on and so on. And they would debate with one another to figure out what the best course of treatment would be.
8 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000719597219