It’s that time every four years when, like many American sports fans, I get passionately interested in soccer. It takes the first week of the World Cup for me to remember all the things I learned about soccer in the last World Cup. Then, most years, I’m able to add a little bit to that store each time. I imagine my soccer knowledge like the rings of a tree. There’s a very thick ring every four years. Then it’s eroded at the edges until another four‑year ring is wrapped around it.

At the core of that tree, the heartwood of my soccer memories, is my own soccer career, which took place largely in the American Youth Soccer Organization. AYSO. The league’s motto is the delightfully egalitarian “Everyone Plays.”

Mostly what I remember about childhood soccer is that we always had orange slices at halftime, and then those mini-sized cans of soda at the end of the game. Every time.

But the real hallmark of little kids’ soccer isn’t the refreshments, it’s the swarm.

The swarm — every child clumped together around the ball, moving as one chaotic, loosely organized mass across the field, stumbling, kicking the ball against one another’s shinguards. Occasionally, the ball squirts free. The fastest children give chase, and they may have a moment in the sun, but they are quickly swallowed up as the swarm reconstitutes itself around them, and order, such as it was, is restored.

This may surprise you, but I think the swarm can tell us something about artificial intelligence.

[

](https://substackcdn.com/image/fetch/$s_!J5vu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58165744-fda3-4987-8994-de3365be94eb_2048x2048.png)

Subscribe now

Salience

The little kids on the soccer field are disregarding their coaches’ commands and swarming the ball because the ball has salience.

The word salience comes to us from the Latin ‘salire’ meaning to leap, which 15th century architects used to describe the protruding parts of a building. It came to mean anything that was prominent or conspicuous, and in the 20th century, cognitive scientists studying attention started using “salience” to describe the quality of attracting attention. Something has salience, or is salient, if it attracts our attention. It leaps into our consciousness.

Even though we often describe external things as “salient,” salience is not an inherent characteristic of a thing. Like beauty, salience is in the eye of the beholder.

One year, my soccer team won our league, and we all got a nice trophy. That trophy was the most salient thing in my room for a few days, catching my eye every time I looked in its direction. Six months later, it was another forgotten trinket gathering dust. Trophies, like anything else, are only salient because we make them so.

Some things are hardwired in our brain to generate salience. The color red for example, tends to attract our attention. We notice when something changes. Movement catches our eye, and loud sounds are more salient than quiet ones.

The soccer ball is salient for those kids in part because it’s a clearly defined white sphere moving unpredictably against green background. It’s always going to catch their eye.

But we aren’t victims of visual contrast. Our brains fine-tune the salience determination with our own preferences. The soccer ball is made even more salient by the fact that kicking it, and then eventually kicking it into a goal, is the point of the game. And kicking things is fun. And everyone else is doing it.

Something that’s impressive about the human brain is that it keeps track of multiple salient things at once, and directs our attention between them. It can even keep track of potentially salient cues, like the sound of the referee’s whistle indicating it’s time for orange slices. Getting through our day is largely a matter of directing our attention. And for the most part, we’re quite good at it.

Kids, not so much. My teammates and I in the soccer swarm knew what position we were supposed to play. In drills, we formed lines and ran down the field keeping our distance. But put us in a game, without that structure, where there’s only one ball and the goals are real, and we swarmed. All those secondary items of interest the coach stressed during practice are swamped by the astounding salience of the game-day soccer ball.

It takes our brains years to build the apparatus of cognitive systems that balance our interest in the ball against our coach’s instructions to stay on our side of the field, to get open for passes, not to swarm. These are the same systems we use to follow the plot of a movie, drive a car through traffic, or listen to a podcast cluttered with silly sound effects.

Artificial brains face the same fundamental challenge. When AI is exposed to a complex input, it has to figure out where it should pay its attention to. It has to assign salience. Here in mid-2026, a few years into the ChatGPT era of AI, we’ve advanced the frontier about to the level of a 9-year old playing AYSO soccer.

Failure Modes

The breakthrough innovation that enabled the modern chatbot is called the “transformer” architecture. In 2017, researchers at Google Deep Mind demonstrated that if you had access to enough computing power, you could get much better results out of a large language model using a relatively simple design and making a vast number of calculations. Regular listeners might remember the pachinko machine from Episode Two, What Is It Like to be a Large Language Model.

One of the interesting things that happens inside a sufficiently large model is that it develops its own internal system for identifying salience. Patterns in the models weights emerge that appear to analyze and react to specific aspects of the user’s input. One might determine the meaning of individual words, another might track grammatical relationships between those words. Researchers calls these patterns “attention heads.” That term always makes me visualize disembodied heads emerging from a sea of matrix-style numbers. Super creepy.

Transformer models are incredibly powerful, but like the swarming kids, they struggle with balancing competing objects of attention. Once they identify something as the bouncing ball, it is difficult to get them to pay attention to anything else. If you like, you can picture this as those disembodied matrix data heads all converging on one stream of symbols. I prefer not picture it that way. It’s not quite accurate and it’s really, really creepy.

AI researchers call repeated similar mistakes by the models “failure modes,” and there are hundreds of papers published every year analyzing the many, many failure modes of AI. These can range from “Can’t spell strawberry” to “Takes instructions too literally and turns the earth and everything on it into a giant paperclip factory.”

Several common failure modes all seem to turn on the model’s limited ability to manage multiple sources of salience.

A pronounced problem that appears in almost any longer conversation is recency bias. Let’s say you start a conversation with a chatbot to plan a dinner party. You decide on things like decorations and activities, and then turn to the menu. It offers you several ideas, and one of them happens to have beets in it. You tell it, great, my friends like beets, let’s do that one. Next thing you know, it’s giving you meal plan featuring beets in every dish, beet themed table decorations, and fun song your guests can sing about beets.

Obviously, this is a hypothetical, beets are disgusting, and no sane person would serve beets to guests they want to return, but you’ve probably seen this pattern of behavior. If you set out multiple goals for a project, whichever one you start with will take over the conversation, and the model will ignore the other goals. If you ask the model to make a small change to a minor point made on page three of a document, it will rewrite the opening paragraph to highlight the change.

The model starts out playing left fullback, like its coach told it to, and for a few minutes it holds its position, but everyone over by the ball is having so much fun, and the ball is just waiting for someone to kick it. It’s so easy to swarm after the bouncing ball of salience.

Share

Ambiguity

For the most part, this sort of swarming behavior can be managed if the human user is careful to frequently remind the model where it is in the larger sea of salience. It’s obvious on the face of the conversation that model is drifting, and so long as the user doesn’t drift along with it (recency bias is a human failure mode as well) the user can nudge it back on course.

But there’s another way in which this swarming tendency can cause problems that are more pernicious. To illustrate this failure mode, let’s move from the suburban soccer fields of childhood to the emerald green pitch of the modern professional game.

If you know just one current soccer player, it’s probably the Argentinian superstar, Lionel Messi. Perhaps the greatest player in history, Messi is known for being a keen observer of the game. He famously spends a lot of the game just walking through the midfield, turning his head back and forth to take in what everyone else is doing. Clearly, he can maintain his attention on many points of salience at once!

For the past few years, Messi has been playing in the United States, but he made his professional reputation with Barcelona, one of the most storied clubs in professional soccer. Between 2008 and 2012, Messi teamed up with a coach, Pep Guardiola, to put together one of the greatest runs in sports history. The team won essentially everything you can win in European club soccer, and Messi was voted the player of the year four years in a row. Guardiola then went on to coach an English team, Manchester City, where he ran off a similar run of success, including 6 Premier League titles in 10 years.

At both Man City and Barcelona, Guardiola was known for tactical innovation - he arranged his players on the pitch in novel ways, and he gave them new instructions for how to attack the opponent. The strategic imperative behind Guardiola’s tactics is to create ambiguity for the other team, and reduce ambiguity for his own players.

Guardiola’s most famous use of ambiguity came when he created a new position for Lionel Messi. In 2009, in a crucial late-season match against arch rival Real Madrid, he had Messi take over the central “striker” role but with a twist. Instead of playing in the classic striker position, at the front of the action, as close to the opponent’s goal as possible, Pep had Messi drift back upfield, into the mix with everyone else.

This put Real Madrid’s central defenders in a bind. If they stayed in position, they were letting Messi roam uncovered in the heart of the pitch, free to receive the ball, make runs, and generally show off his otherworldly skills. But if they followed Messi upfield, they were leaving their critical area, in front of their own goal, unguarded for Messi’s teammates to move into. As Messi’s teammate Thierry Henry put it, when faced with this choice, either “you die, or you die.”

Barcelona went on to win that match 6-2, Messi had two goals, and Henry had two as well. The soccer world fell in love with the so-called “false 9” - a striker who lurks in the middle of the pitch, rather than playing right up to the opponent’s goal.

Forcing defenders to make impossible choices, while giving your own players easy and obvious ones, is hardly a new idea. Even the false 9 itself wasn’t new, it had been around at least since the 1950s, but Guardiola revived it at just the right time with just the right player.

Manufacturing ambiguity is all over sports. In American Football offensive coordinators develop elaborate schemes to try and force defensive players into ambiguous situations, where they aren’t sure who they should be covering. Forty-Niners running back Christian McCaffrey is one of the most valuable players in the league because he has the skill set to play two positions, running back or wide receiver, and so just by stepping on the field, he creates ambiguity.

In basketball, a new generation of 7 footers who can shoot from long range puts defenders in the same bind that Guardiola and Messi placed Real Madrid. Defenders can stay in position under the basket, but then the big man will have an easy three point shot. If they follow him out there, they leave the basket unguarded.

Language is littered with ambiguity. Dissect any passage of human prose, and you are likely to find multiple possible meanings. In the last example, “shoot from long range” obviously meant taking a basketball shot, not firing a weapon, but you could only determine that from the larger context.

But some ambiguity can’t be easily resolved. If you say, “Can you bring the drinks for the soccer thing today?” then I’m not sure if you mean the half-size soda cans for the kids game at the park, or the half-sized keg that our Scottish friends will drain by halftime of our World Cup watch party.

AI models, however, blow right past these genuinely ambiguous situations. They pick an interpretation, often one that seems far from obvious to a human, and swarm it. Once it’s decided you wanted a keg and not cokes, it can difficult to convince an AI there was ever any other option.

What’s interesting is that when asked, models are actually quite good at identifying ambiguity. Researchers call ambiguous queries “Under Specified” which I think is rather funny, it suggests it’s our fault for not being good enough at writing queries. Anyway, when you give a model an intentionally “under-specified” query, and ask it, “Is this ambiguous?” models will identify the ambiguities with great accuracy. But if you just give the model the under-specified prompt directly, the model typically doesn’t say, “Hey this is ambiguous, I’m not sure if I should cover Messi or stay in my area.” It just picks one, and runs with it.

Humans, when faced ambiguity like this, have a powerful tool. We ask questions. “Which soccer thing?” likely resolves our uncertainty about bringing a keg to the party or sodas to the kid’s game, and defensive players facing a novel offensive scheme can ask their coach what do. The coach might not have a good answer, but at least the players have put the decision on someone else.

AI models, however, are terrible at asking questions. Think about the Deep Research tools most of the chatbots have now. They ask 3-4 of the same standard questions of every request (“Do you want me to look at just the U.S. or consider global implications?” “Should I limit this research to primary sources or include well-vetted secondary sources as well?”) and charge ahead.

Sometimes, we’re ambiguous on purpose. We’re unsure about something, and we want to convey that. I might have a theory why my code isn’t working, or why the middle of a speech I’m writing drags, and I want AI to assess that theory. A good way NOT to get a good response is to submit something like, “I’m not sure about this but I think,” and then describe the cause you’ve hypothesized. More often than not, the model will ignore the ambiguous part of your query, and take your prompt as a concrete instruction, changing the code or rewriting the passage as if certain about the problem.

Man-Marking

There’s no single answer to the challenge posed by the False-9, and when that 9 is an elite dribbler and passer as well as a threat to score, it’s terribly difficult to defend. Besides Messi himself, England’s Harry Kane has been using the freedom of the position to cause problems for defenses in this World Cup.

But a popular response, and one that has the advantage of reducing ambiguity for the defense, is assigning a single defender to shadow the false-9 wherever they go. This is known as marking or man-marking. It weakens the structure of the defense, because the remaining players must still divide up the entire field, but it takes the problem of how to handle the roaming striker off the table. Sometimes, teams will convert to a mostly or fully man-marking system - what in basketball and American football is known as man-to-man. By the end of Guardiola’s career at Manchester City, he faced man-marking schemes in nearly every match. It’s much, much simpler for the defense, but also more physically demanding - defenders will find themselves on an island against a skilled player with the ball, and getting beat just a couple times a game can be enough to lose.

At the risk of stretching this extended soccer analogy past its breaking point, man-marking is not a bad way to think about how models handle salience and ambiguity. Rather than maintain a structure, with some players guarding empty space, and dealing with the ambiguity created by crafty managers like Guardiola, AI models just grab the first player they see and stick to them.

I have a highly anthropomorphized theory about all of this. So much effort has been put into salience, that the models fear and loathe ambiguity. It’s as if we trained them up to play striker, but every so often, we ask them to play keeper. “User your hands,” we say, “that’s allowed now.” “But, keep the ball out of the net, don’t let it in.”

Maybe we’re right on schedule though. The transformer model turned 9 years old in June, and nine is about the age at which kids develop the skills to hold their position and stop swarming. It takes nature 9 years just to get the basic salience tools in place, then it can start handling ambiguity. So, happy birthday large language models, have an orange slice.