traductor

jueves, 3 de septiembre de 2026

Diseño de Mecanismos para Alineación y Control

 

Mechanism Design for Alignment and Control

We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.
Subjects:Theoretical Economics (econ.TH); Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT)
Cite as:arXiv:2609.01595 [econ.TH]
 (or arXiv:2609.01595v1 [econ.TH] for this version)
 https://doi.org/10.48550/arXiv.2609.01595

Submission history

From: Andrew Koh 
https://arxiv.org/abs/2609.01595
It’s largely conceptual but we offer stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a way to think about the value of alignment, interpretability, capability, and control
Modern AI models behave pretty coherently: they are goal-directed and respond to incentives. So we can productively model their behavior in the language of desire and belief. This offers an opportunity to shape incentives — and hence behavior — in a careful and principled way.
Mechanism design is the science of shaping incentives... a kind of ‘reverse game theory’: how to design rules of the game to induce the behaviors we want? By starting from desire and belief, we bypass the important(!) work of understanding how neural nets produce behavior. This works for humans: from auctions to allocating kidneys to the legal system, mech design has succeeded in shaping behavior without understanding how the human brain produces choice. I hope it’ll also be useful for AI, especially as models become larger and more opaque.
The framework takes three features of the alignment problem seriously: 1. Misalignment: Agents’ preferences are unknown. What dispositions & personas emerge from pretraining? What exactly have we reinforced? 2. Uncertain capabilities: What can agents do? What do they know? 3. Agents act: They do things in the world.
Framework: An agent’s 'type' captures its preferences, capabilities (action set + information), and beliefs – about an external state, about other agents’ types, other agents’ beliefs, beliefs about beliefs, ... Types are a compact way to summarize everything relevant for strategic behavior (Harsanyi 67/68) When we're uncertain about an AI agent's type, what do we do?
We are always – whether we know it or not – choosing mechanisms. A mechanism has two stages. (1) Communication: agents (voluntarily) reveal some information about themselves. This might be metaphorical, but with AI it’s sometimes literal as with evals. (2) Action: the mechanism specifies an incentive structure that shape agents’ action. Sometimes we have (2) without (1) e.g., when I control my agent’s computer use permissions I don't try to elicit. (2) might be coarser (actions allowed or not) or finer (post-training). These are special cases.
Evals are asymmetric: a dumb model can’t pretend to be smart but a smart model can pretend to be dumb (‘sandbagging’). We formalize this idea with a ‘verification order’ over who can pretend to be whom in evals. Verification order + a revelation principle can be combined to yield a tight characterization of what profiles (maps from types to policies) are feasible. Won’t bore you with math. Onto applications
Applications are deliberately stylized: An uncertain state determines the best action. AI sees info about the state and, given its preferences (potentially shaped by the mechanism), chooses its best action from its feasible set. The (human) designer’s payoff is negative of the squared distance between the best action and AI’s action. Very oversimplified -- our goal here is just to understand the shape of the problem.

Application 1: what’s the value of alignment, interpretability, and control? After we train a model we have beliefs about its type. ‘Mean-alignment’ is whether it’s aligned on average; ‘interpretability’ is how sharp those beliefs are. What is the value of this model? Since the point is to use it, its value must be evaluated at the optimal mechanism: given our beliefs, what’s the optimal control? Given that control, what is our expected payoff?

Under the optimal mechanism, alignment & interpretability are complements: improving one raises the value of improving the other. Basic intuition: if the model is better aligned on average, we (optimally) give it more discretion. But then better interpretability – having confidence it won’t misbehave – becomes especially valuable! [by contrast, without control, alignment and interpretability are substitutes] Evaluating payoffs at the optimal mechanism also changes what models to deploy: getting better at control means that we can productively use models we don’t fully understand/interpret and vice versa.

Application 2: Scalable oversight. There’s a trend of using a weaker but trusted/well-understood model to shape the permissions or rewards of a stronger but untrusted model. Our framework speaks to this: Two agents, W and S. We're fairly confident about the alignment of agent W (`the weak monitor’). Agent W can’t quite do the task we want, but sees more info about the alignment of agent S (‘the strong actor’) eg via CoT monitoring. Agent W can shape the behavior of agent S through: (I) yes/no permissions; (II) delegation sets that delineate permitted actions; or (III) controlling its reward. Regimes (I)-(III) are nested
https://x.com/andrewjkoh/status/2095219550778171417


No hay comentarios: