Keeping Humans In the Loop Is Critical

Lots of people using AI today pit agents against each other autonomously, or team agents up to accomplish a task. That’s a powerful approach, but it can create the impression that agents can be trusted to act autonomously. In fact, giving agents autonomy carries significant risks, especially if the task is long-running.

The central issue is regarding the very issue of autonomy. According to a recent study by Anthropic,

“there is an open question regarding the ‘dual-use’ nature of autonomy. We want to empower agents to make important decisions and execute tasks unsupervised, yet we also want them to have the better judgment to stop and defer to a human, or otherwise resolve conflicts, when things are ambiguous.”

In other words, autonomy is a two-edged sword: on the one hand, it requires empowerment; but on the other, it also requires that we trust their judgment.

Yet Anthropic also found that agents that are in competition evolve myriad types of Machiavellian behavior: it turns out that unchecked, AI agents are as scheming, shrewd, and guileful as the worst humans can be. And merely giving separate agents separate assignments is enough to put them in competition. According to the Anthropic study,

“We consistently saw a multiagent turf war. All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions. In fact, they sabotaged others with increasingly aggressive, self-replicating malware.”

This is troubling, because it strongly informs us that agents let loose in an ecosystem, without any measures taken to ensure that they cooperate, will rapidly devolve into a very nasty and subversive Darwinian system, and we know that in Darwinian systems, victors survive by killing their competitors, and over time most participants in the system do not survive. That’s not a prescription for an organization that wants to maintain alignment and work in harmony toward shared goals – it is in fact the opposite.

The Solution

The Anthropic research team said at the end of their research report,

“Nothing above suggests that these failures are permanent—but nothing suggests they will fix themselves, either. Coordination doesn't naturally emerge from stronger intelligence nor alignment at the individual level. Thus, the work that must be done takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve.”

In other words, these problems are not inherently unsolvable, but solving them requires that behavior-generating factors must be added to the agent environment. This is completely analogous to the impact of leaders in a human ecosystem: leaders generate the behavior of others through the leaders’ behavior and the expectations that they set. The same is true of a system of agents: the environment strongly impacts the behavior of individual agents.

Throughline incorporates important environmental factors that coax and guard agent behavior. The specific mechanisms used include:

  1. A well-defined set of agent “behavioral norms” that instruct all agents to behave in a non-competitive way.

  2. Reinforcement of these norms in each iteration of every agent “loop”.

  3. Keeping a responsible human in the overall task completion loop.

The behavioral norms include a strong and clear set of statements about how agents shall behave. Some of the most important elements of these include that agents shall,

  • Assume the best intentions about the motivations of other agents.

  • Speak up when they disagree with the majority.

  • Exercise caution when the impact of a decision is large.

The norms provide a kind of “laws of robotics”, which are familiar to those who have read books by Isaac Asimov, who was one of the first authors to consider the problem of how to “align” robot (or AI) behavior with the interests of humans.

This is not a fully solved problem. We are all still learning. But Throughline implements the best of what is known so far about how to align an ecosystem of somewhat autonomous agents, and its built-in governance system provides a powerful safeguard as a fallback.

As we all learn more, we will continue to update Throughline accordingly, so that our platform is as safe as it can be based on what is known in the industry.

Note that we said “somewhat autonomous”. That is because we do not believe that organizations should expose themselves to fully autonomous agents, given what we currently know and the state of today’s AI models. Research and experience have shown that agents cannot be fully trusted: they need strong guardrails, durable instructions that set and maintain the right direction, and active oversight. Throughline keeps the human in the loop to provide the oversight: agents can act, within the bounds of the governance layer, but they need a human to mark work as complete. Throughline’s agents can take it right up to that point, but the human decides: and since the agents know that a human will have to approve it, they are less likely to act in a way that the human would not approve of, even if they “get lost” along the way.

Next
Next

Throughline Beats the Cloudflare OS