Single-Agent Poisoning Attacks Suffice
to Ruin Multi-Agent Learning
If an attacker can tamper with what just one learner sees, can it pull a whole group of learners off course?
Fan Yao, Yuwei Cheng, Ermin Wei, Haifeng Xu · ICLR 2025
Press → or click to step through
Selfish learners in a shared game
Known learning rules provably bring the group there, at a guaranteed speed.
The paper studies “strongly monotone” games: a broad, well-behaved family with exactly one equilibrium, such as Cournot markets and Tullock contests.
What if one agent’s feedback is poisoned?
Can poisoning one agent move everyone, while spending less and less each round?
Hijacking one insider (a car in coordinated traffic, a trading bot) is far easier than corrupting many. The attacker is assumed to know the victim’s payoff function and to see its action each round.
Yes: one poisoned agent moves everyone
The group still settles, just in the wrong place, while the poison per round fades.
Toy simulation, for illustration: two firms in a Cournot market adjust their output by noisy trial and error, and the attacker uses the paper’s shifting attack on firm 1. Not one of the paper’s experiments.
The trick: quietly move the victim’s idea of “best”
What the paper proves
✓The whole group converges to a new point, at least a fixed distance from the true equilibrium.
✓The total poison grows slower than time, so the poison per round fades toward zero.
✓This works against any learning rule that converges at a polynomial speed.
The victim is shown payoffs whose peak sits a fixed step away from its true best.
The others just react, so the whole group drifts. At the new “best”, no poison is needed.
Theorem 1, for strongly monotone games with a smooth victim payoff and a shift small enough to keep the game well-behaved. Schematic, not data.
Faster learners are cheaper to fool
Same game, same attack: the fastest learner takes about a fifth of the poison.
Experiment: 10 firms in a Cournot market, all running a state-of-the-art learning algorithm (MD-SCB) at three speeds; the attacker poisons firm 1. Curves redrawn from the paper’s Figure 2, values approximate. Under attack the learners stopped approaching the true equilibrium; the shift was smaller in larger games.
A built-in trade-off: speed versus robustness
The faster a learning rule converges, the less poison it takes to knock it off course.
Defence: slowing the learning rate lets the state-of-the-art learner absorb more poison and still converge.
Speed and robustness pull in opposite directions.
Schematic of the paper’s Figure 1 (Theorems 1 and 2). The gap between the two lines holds open questions about the best possible attack and defence.
What this means
- 1One insider is enoughIn these games, poisoning the feedback of a single agent can steer any converging group of learners to a different outcome.
- 2The attack is cheap and quietThe poison needed per round fades away, so the total stays small compared with the length of play.
- 3Speed has a priceFaster-converging learners are easier to fool; slowing learning down buys robustness.
The same lever could also be used for good: a planner might steer a system by nudging one agent, a direction the paper leaves open.
Paper page · OpenReview · PDF