ChatGPT Just Became a Work Agent artwork

ChatGPT Just Became a Work Agent

The AI Daily Brief: Artificial Intelligence News and Analysis

July 10, 2026

OpenAI’s new ChatGPT Work brings the agentic systems that transformed coding into the broader world of knowledge work, allowing AI to operate across apps, files, and long-running projects. NLW breaks down what the new harness means, how GPT-5.
Speakers: Nathaniel Whittemore
**Nathaniel Whittemore** (0:00)
Today on The AI Daily Brief, more new models plus a big harness update from OpenAI. And before that in the headlines, Cursor also appears to be developing a harness to go after the larger knowledge work sector. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Rackspace, Blitzi, and Airtable. To get an ad-free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. And of course, to learn more about sponsoring the show, head on over to aidailybrief.ai/sponsors, or send us a note at sponsors at aidailybrief.ai.
Quite appropriately, given that our main episode is about a big harness update, we kick off our headlines with news that Cursor is planning a general purpose agent to compete with Claude Cowork. The information reports that work began on the project in April, shortly after Cursor signed their deal with SpaceX. The agent is expected to use Grok 4.5 and will be Cursor's first project aimed at anyone other than professional coders. Called Sand, it's designed to function as a personal assistant performing standard office tasks like dealing with email or working with spreadsheets. It sounds like it could also eventually become a unified platform, with the information's reporting suggesting that the agent will also be functional at AI coding. Sources said the platform was rolled out internally in June, however, it's still unclear whether it will get the green light for a public release or when that would happen. Certainly the product suggests that Cursor, now part of SpaceX AI, is looking to grow beyond their traditional user base of software engineers. This makes sense, as we'll see a huge theme of all of the announcements today with OpenAI are all about taking what has been working in coding and bringing it to a broader set of knowledge work.
Now speaking of what's working with coding and what's not working, OpenAI has captured the zeitgeist and declared that the leading coding benchmark is bunk. In a new report, OpenAI audited SweetBench Pro and found the benchmark to be sorely lacking. In their testing, they found that 30 percent of the tasks on the benchmark were broken and are now formally retracting their support of the benchmark. Many of the issues stem from some of the tasks being public, which can distort results by having those specific problems be included in training data. Others had hidden requirements, contradictory instructions, overly strict tests, or incomplete grading criteria. Their conclusion was that SuiteBench Pro no longer reliably measures frontier coding capability. To be clear, the shift away from SuiteBench was already well underway. Cursor has been using their own proprietary benchmark for months, while Cognition and Databricks also launched their own benchmarks this week. I think we're officially at the point where there's going to be lots of introduction of new benchmarks, companies are going to present all of them, including I guarantee they will still present SuiteBench Pro, and mostly just wait for people to have the vibes that confirm whether the benchmarks are bunk or not.
Now, staying on OpenAI for a minute, but going to a very different area, the company has published a new statement that explains their approach to government and military partnerships. Announcing their new national security principles, OpenAI writes, we believe democratic society should be able to use AI to protect people, defend critical infrastructure, deliver public services, and respond to emerging threats. But increasingly capable AI systems must be deployed in ways that reinforce democratic accountability, meaningful human judgment and the rule of law, and strengthen democratic institutions rather than concentrate power. Distilling the principles into a few key points, OpenAI explicitly stated that they will not support the use of their technology for mass domestic surveillance, high-stakes decisions including decisions over the use of force without appropriate human judgment and accountability, or uses that evade legal obligations, oversight, and accountability. Now these you might recognize are essentially the same as Anthropics Red Lines, and given that that created such chaos with the government, it's not clear what the goal of this document is supposed to be. But I guess at least we have a clear articulation of what they stand for and something to build from in their conversations with the government.
Speaking of the overlap between the government and AI, Anthropic has appointed former Fed Chair Ben Bernanke to the board of their Long-Term Benefit Trust. Bernanke, of course, was appointed Fed Chair by the W. Bush administration in 2006 and served until 2014, presiding over the events that led up to and followed the global financial crisis. He is one of the more controversial Fed Chairs in recent history, with views on him differing depending on whether you think the bank bailouts in 2008 saved the world from a depression or you view them as the original sin that accelerated the wealth divide in the United States. When Claude was asked to provide an objective viewpoint on Bernanke, it responded that he's generally well regarded but has critics on both the left and the right who view him as, quote, emblematic of an unaccountable technocracy protecting elite interests. So what then is Anthropic's Long-Term Benefit Trust? In their own words, they write, Anthropic is a public benefit corporation, meaning the company was created to balance commercial success with generating social and public good. The Benefit Trust exists to help the company responsibly maintain that balance over the long term, providing a check on how Anthropic develops and deploys AI. Essentially, this is an independent oversight and advisory board that sits on top of Anthropic's normal board of directors. Importantly, the Benefit Trust has the power to elect or remove one member of the corporate board, which escalates to two and then three according to time and funding-based milestones. By next year, they will have majority control over the corporate board, albeit with a shareholder override that requires a supermajority vote against the actions of the trust. Unlike normal corporate boards, no one on the board of the trust is allowed to be a shareholder.

24 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000776302081