**SPEAKER_1** (0:00)
This podcast is sponsored by Google. Hey folks, I'm Amar, Product and Design Lead at Google DeepMind. We just launched a revamped vibe coding experience in AI Studio that lets you mix and match AI capabilities to turn your ideas into reality faster than ever. Just describe your app and Gemini will automatically wire up the right models and APIs for you. And if you need a spark, hit I'm feeling lucky and we'll help you get started. Head to ai.studio slash build to create your first app.
**Nathan Labenz** (0:31)
Hello, and welcome back to The Cognitive Revolution. While we often discuss sovereign AI in the Silicon Valley AI bubble, we rarely hear directly from the technical leaders who are actually leading national AI projects. And so today, I'm very glad to share my conversation with Marek Kozlowski, who's leading project PLUME, which stands for Polish Large Language Models, in his role as head of the AI Lab at the National Information Processing Institute of Poland. Poland, with a population of 38 million and GDP of roughly 1 trillion, roughly 10% and 3% of the United States, respectively, is an interesting and in some ways a representative case study. It clearly doesn't have the resources required to compete with the US and China at the AI frontier, but it does have strong technical talent, a real sense of pride in its language and culture, and a deep desire to control its own technological destiny and avoid domination by global superpowers. So what does that mean in practice? As you'll hear, Marek's strategy relies on the core belief that by training small models for a particular local language and cultural context, countries like Poland and projects like PLUME can compete with the latest frontier models, all while retaining control, preserving data privacy, and achieving a major cost advantage. In this conversation, we dig into the strategic realities that motivate projects like PLUME and the technical challenges they have to overcome to succeed, including how today's frontier models, which are trained on overwhelmingly English and Chinese data, fall short in other languages. Why this problem is actually getting worse from one generation to the next, as frontier model developers prioritize things like coding performance above support for niche languages. How EU regulation prevents European AI Builders from conducting massive web scrapes and instead forces them to rely on more focused data curation projects. How the Polish government is thinking about investing its finite resources across data, compute, and talent. The language adaptation techniques that Marek's team layers on top of Lama and Mistral base models so as to inject local knowledge without needing to start from scratch. Why they haven't yet had to worry about developing a constitution or other explicit articulation of values for Polish AI systems. And why government agencies and national champion companies are often better served by smaller models, fine-tuned for specific tasks, and served locally than by massive generalist models served from the cloud. Overall, Marek's mix of realism about the challenges of competing with global leaders and his positive vision for transparently created, locally controlled AI is a great window into what AI leaders around the world are thinking and doing to maintain AI sovereignty. So with that, I hope you enjoy this deep dive into the meaning and training of Polish AI with Marek Kozlowski. Marek Kozlowski, Head of the AI Lab at the National Information Processing Institute of Poland. Welcome to The Cognitive Revolution.
I'm excited for this conversation too. We met not too long ago at an AI event in Las Vegas, the Enterprise Technology Leadership Summit, and I thought it was really interesting to double click on everything that you're doing because in the United States and in the sort of Silicon Valley AI circles that I spend most of my time in, there is this ongoing conversation about sovereign AI. And I think it's funny that a lot of this conversation happens in the Silicon Valley bubble and sort of makes a bunch of assumptions about what other countries feel the need to have, you know, are aspired to create, you know, what's driving those decisions. And I don't hear too much from primary sources of people that are actually doing the sovereign AI projects around the world. So I was excited to meet you and learn more about what it is that you're doing in Poland. Poland, obviously, I think, obviously to me, you know, is a country with a lot of technical skill and, you know, very distinct culture, obviously its own language, proud tradition. And so I'm really interested to get into it and figure out what sovereign AI means in the context of Poland.
**Marek Kozlowski** (4:41)
Yeah, once again, thank you for the introduction and for introducing my person and showing the idea. The idea is that I called it slightly broader, not only the sovereignty, but also the creating the localized LLMs. Yeah, because the localized can be the national LLMs, but also the domain-oriented LLMs. Yeah, and I create the, maybe not I create the idea, but I am promoting the idea of the localized LLMs. It means the LLMs adapted to the language or domain, because they can be also a domain, and they are in this domain or the language, they have higher quality understanding, texts in this language or domain, and have the higher quality and they are able to create the higher quality texts in a generation step. It means that building the localized LLMs, they can be of course adapted to the language or domain, has two goals. First of all, to improve the understanding in this domain or the language, but also give the possibility to generate the higher quality texts, of course, in the aspects like the linguistic and cultural aspects.
71 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000739976444