**SPEAKER_1** (0:00)
If you want to get access to this episode and my next 30 episodes, all ad free, so there'll be no ads on them, go check out my podcast, AI Chat. You can go search for that on Spotify or Apple. It's AI Chat. I'm gonna post all of these news episodes, and I'm also posting interviews, like I just interviewed the CEO of Cohere. They've raised over a billion dollars for their AI model, talking about what they're gonna be spending the money on and the direction of the AI industry, along with all of this new stuff. So if you wanna go check it out with no ads for free, it is AI Chat. Thinking Machines by Miriam Marotti, who famously left OpenAI after she was the CEO when Sam Altman got kicked out, have released their first open weight AI model called Inkling. OpenAI has built something called Gpt-Red. It's an LLM super hacker, and they built this to harden their own models. AI Box has shipped an MPC server that brings 80 different AI models into Cloud, ChatGPT and Gemini. AWS is committing $1 billion to a forward deployed engineering organization. OpenAI and Anthropic are both scaling theirs. Apple Intelligence has been cleared for a China launch with Alibaba's Quen, and Meta plans to sell and resell their AI compute following what SpaceX is doing, and kind of copying their playbook. Let's kick this off with Thinking Machines. This is a company that I'm really rooting for. They've raised over a billion dollars, but they haven't come out with anything super groundbreaking. But what I will say when it comes to a lot of these models, they're so expensive, and they take so much time that when you think of something like even Anthropic, it felt like OpenAI had just run away from them, and they were never going to catch up. And a little by the little, they pick their lane, they make their product better, and they're able to get market share until the point where Anthropic has now exploded and is making more money than OpenAI. It surpassed them in revenue. I think we might see a lot of that same strategy played out by a lot of different AI companies if they don't kind of shrivel up and die, or get sold off for parts, or an aqua hire or something like that. And so Thinking Machines is one of these companies that I'm really excited about, and I hope that they really go places. But this is run by Miriam Marotti, who is famously the CEO of OpenAI when Sam Allman left. And she has created this thing called Inkly. It's an open weight AI model. It has 975 billion parameters. And companies can download this and customize it themselves. So you don't have to go and pay for API access, OpenAI or Anthropic. You can actually just go and fine tune this all on your own. You can fine tune on your own data. And the bet that they want people to take is that that is going to be better than a one size fits all model that OpenAI or Anthropic or Gemini or any of these other big labs are selling. Inkly was trained on 45 trillion tokens of text, image and audio and video. They did it in about nine months, which is way faster than OpenAI's five year timeline or Anthropic's three years. In a test with Bridgewater Associates, the financial model trained on the hedge fund's own expertise scored 84.7% on financial reasoning benchmarks at roughly 1 14th, the cost of the top models from Anthropic and OpenAI. We're seeing some massive improvements when you're taking these models and you're fine tuning it on your own data.
The model activates only 41 billion of its 975 billion parameters per task, and users can dial up thinking efforts to trade speed for accuracy. So I'm rooting for them, but time will tell how well this model does. We just learned that OpenAI built something called Gpt-Red. This is an AI model trained to attack their own systems, and they use it to catch security flaws before release. So the model cut successful attacks on Gpt 5.6 from over 90% down to 23%, which basically makes it one of OpenAI's most secure releases yet. One particularly interesting attack that it was able to discover is called a Novel Fake Chain of Thought, which basically is tricking a model into making up fake reasoning steps. So it's kind of similar to convincing someone that, you know, one plus one equals three, and then, you know, saying that I already checked the math, that's what equals, and if it equals that, then therefore, and you, you know, go trick them on the next thing. So they found that when they gave the same task as a human Red teamer who tested GPT-5 in 2025, Gpt-Red found more effective attacks than the humans. They also got it to go and hack like this third party vending machine agent, which is kind of funny. It does have a bunch of limitations, so it's not very good at back and forth conversation attacks, and it's not very good at exploiting images that are embedded into malicious text. The idea behind this is actually functionally working, is that Gpt-Red is automating how all of the security holes are found, because it puts an attacker model against a defender model, and they put them in these kind of simulated real world environments, and they're battling it out. OpenAI is not releasing this externally, so they're not giving this to other people to test, which is interesting, right? Because we had the whole Anthropic Mythos model, which was really good at security exploits, and they gave it out to all of the top labs and said, hey, like go harden all your security with this. So OpenAI is not giving this out externally beyond just using it for themselves. And I mean, it's got a huge massive drop in the success rate of a lot of these exploits. I think it shows AI-powered red teaming can be just as effective or more effective than humans. I mean, you can just brute force way more tests than humans, and it's trained off of what humans are doing. But at this point, it's doing better than what a lot of the humans are doing. AI Box, which full disclosure is my own startup, has just released an MCP server, and we are allowing people to plug 80 different AI models straight into Cloud, ChatGPT, Gemini or Cursor. Basically it lets you use any model inside of whatever assistant you already are used to. Personally, I use Cloud all day long, although I'm switching over to ChatGPT with their new ChatGPT app. It's really amazing. But either way, you get the AI Box MCP, and it allows you to access all of the images that OpenAI can generate. It allows you, if you're on Cloud or ChatGPT, to generate all of the videos that Google V03 can make. And no matter what platform you're using, you can access what 11 Labs can do with audio. I spent basically the entire day today getting ready for a big Facebook campaign that we're launching here at AI Box. And in the past, doing Facebook campaigns usually meant hiring a Facebook ads team. It meant creating tons of creative, which just takes a lot of time.
11 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777145496