**Nathaniel Whittemore** (0:00)
Today on the AI Daily Brief, how to get the most out of frontier models. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzi, Retool, and Airtable. To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at aidailybrief.ai. And a quick note about today's episode, this was recorded in advance. In fact, I am recording it on Thursday afternoon, as everyone freaks out about KimmyK3. So in the meantime, if Dario and Sam have lost their minds and released new versions of Fable and GPT in response to the threat that's tearing value off the NASDAQ, you'll know why I am not talking about it right now. Just a quick little bit of travel on Monday. I will be back on Tuesday with a normal episode. But still, regardless of what is going on in the wide world of models out there, today's episode is all about how to get the most out of the most advanced and newest models. So without any further ado, let's dive in.
It has now been a couple of weeks with this new class of models in Fable 5 and GPT-56.
Now, weirdly, these models have actually been around a little longer than a couple of weeks. There was a particularly long early access period for GPT-56, and Fable 5 was here for a couple days before going away. But at this point now, pretty much everyone has now had these models for some time. In fact, our access to them keeps getting extended and reset. And along with that, people have started to publish their tips and tricks for getting the most out of them.
Now, it is always the case that new models demand new ways of interacting with those models to get the most out of them. But this is exactly the sort of information that can't really be captured in anything like benchmarks and just has to go be experienced through trial and error. As you will see, there are a number of common threads that cut across both 5.6 Sol and Fable 5 that suggest, I think, not just some new ways to get the most out of these models, but for some new patterns of interaction that are going to become increasingly common from here on out.
Now, we're gonna start with some official sources and commentary. Codex team member Eric Provenchar wrote, With 5.6 Sol, a lot of people are still prompting the model exactly as they did 5.5. It's important to note that 5.6 Sol is a lot more tenacious and thorough than previous models. Eric published a prompting guide on the official learn.chatgpt.com site, and while in this case, it is not presented as a way to get more out of 5.6 as opposed to 5.5, there are a few things that are specifically worth noting. One piece of guidance that comes from that tenacity is around setting boundaries. Boundaries, writes the guide, are the few instructions ChatGPT needs to avoid creating extra work or taking an action you didn't intend. Add one when changing the wrong detail would make the result unusable, or when you want to review something before it affects other people. Examples. Keep the approved dates and budget figures unchanged. Use only the supplied sources. Flag missing information instead of guessing. Keep recommendations within the stated budget. Prepare the message as a draft, don't send it. You can see how in each of these cases, a more tenacious model, to use Eric's word, might assert a little bit more agency than the user would be comfortable with and actually go off and do something that ended up not being optimal for whatever the prompter was trying to achieve. One example of these boundaries was being explicit about where you wanted it to stop in actions you didn't want it to take, i.e. don't send the message, just prepare it as a draft. Another example around the approved dates and budget figures was to limit and focus where the tenacious model was applying its attention. When it comes to the injunction to use only the supplied sources, a tenacious model that has access to the entire internet could go off and get lost in some serious rabbit holes if not told not to do that. Now obviously setting boundaries is nothing new, but the point here is that the more powerful the model, the more significant those boundaries become and the more real world consequences there can be if those boundaries aren't set. On the lower end of the spectrum of consequence, there's just burning through way more tokens than you actually needed to, which in an increasingly cost conscious era isn't nothing. But then of course there's much bigger consequences, like setting a message that hasn't been approved yet to a critical customer partner or colleague.
23 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777598531