**Stephen Wolfram** (0:05)
Hello, everyone. Well, I'm gonna go through a piece that I just posted about a week ago. It's called Games Between Programs, The Rulology of Competition.
The basic setup. Whether one's dealing with biology, economics, politics, or a host of other fields, it's common to encounter situations that can be modeled as involving two agents that repeatedly compete with each other. One imagines that at each step, each agent can take one of a certain set of actions, and that then in a kind of classic game theory way, each agent or player gets a certain fixed payoff based on the action they and their opponent take. But how do the agents decide what action to take? We imagine that each agent has a certain fixed procedure or strategy for making its decisions. We imagine that the input to each of those decisions is the sequence of past actions that the agent and its opponent have taken. So there's been lots of work done over the course of nearly a century on particular choices of strategies. But something I've long been curious about is what happens if one systematically considers all possible strategies. And if we think of strategies as programs, this becomes a question to which we can immediately apply ruleological methods. Which is what I'm going to talk about here. Well, to be more specific about the setup, let's assume that at each step, each agent takes one of two possible actions, which we can indicate by red and green. And let's, so let me show you that. Here's the setup. Two possible actions for each player indicated by red and green. I'm going to take the payoffs to be in this, then what we're starting with here, to be the ones for a sort of classic match or not, or so-called matching pennies game, in which player one has the bigger payoff when there's a match between the actions of the players, and player two gets the bigger payoff when there isn't a match. So two greens, player one gets payoff plus one, a green and red, player two, agent two gets payoff plus one, agent one gets payoff minus one. So that's the set up.
So what happens when agents repeatedly play this game? Well, depends on the strategies. So here are a few examples of different choices for each agent's strategy. And so what's happening here is the choice of which action to take red or green is determined by the strategy that each player is using and that it's determining that color based on previous actions that the other player used in this game. So, okay, with this setup, we can kind of plot the cumulative payoffs for the two agents in each of these five different scenarios.
So we can consider the winning agent to be the one that has the numerically largest cumulative payoff, i.e. the one which is eventually on top in these plots after a certain number of steps. And with a criterion like that, we'll be able to rank different programs against each other, and in general, explore the ruleology of competition. So with the basic setup we're using, we can represent all possible sequences of actions by a multi-way graph.
So for any given sequence of actions, there is then a cumulative payoff for each agent, for the particular match or not game. So that's what's shown here is those payoffs for each of the possible paths, each of the possible sequences of actions, we get a certain payoff indicated in this way. So if each agent adopts a particular strategy, this will define a particular path through this multi-way graph. So for the strategies that we just talked about a couple of moments ago here, the paths through the multi-way graph are these ones here.
So what does it take to have a winning strategy? And what we're going to do, we'll consider strategies that are based on several different types of programs. But one basic question we can always ask is whether what turn out to be the winning strategies tend to be based on programs that are more complicated or less show, so, or show behavior that is more complicated or less, or less so. So in other words, if you want to win, should you typically be trying to build up something complicated, or should you instead inspect to be able to find some sort of simple hack that will crack the game, at least usually, let you win? In effect, we're asking whether competition leads to complexity or simplicity. So I've recently looked quite a bit at normal models about biological evolution and machine learning, in which one is adaptively evolving programs in order to maximize some externally imposed fitness function. What I've found is that even when the fitness function one uses is simple, the behavior of the programs that maximize it is normally quite complex. In other words, adaptive evolution will tend to make even a simple, fixed objective be achieved in a complicated way. So, what if instead of having a fixed externally imposed objective, our goal is just broadly to win against other agents? Does such potentially open-ended competition lead us to more complex behavior, or more complex programs, or not? That's the kind of question we're going to be able to explore here by looking at the Rulology of Competition.
52 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773948526