By NHI Mgmt Group Editorial TeamDomain: AI SecuritySource: PenteraPublished August 17, 2026

TL;DR: Even perfect reasoning, prediction, and learning do not solve the value step in AI alignment, according to Pentera, because Arrow’s impossibility theorem and Gibbard-Satterthwaite show preference aggregation cannot be made simultaneously fair, universal, non-dictatorial, and strategy-proof. The real constraint is governance, not computation: humans must choose which axiom to relax.


At a glance

What this is: This is Pentera’s final chapter in a series on AI limits, arguing that value aggregation in alignment is mathematically constrained by Arrow’s impossibility theorem and the Gibbard-Satterthwaite theorem.

Why it matters: It matters to IAM practitioners and governance leads because AI systems that shape access, decisions, and policy cannot be aligned by optimisation alone when the underlying objective is plural, contested, and strategically manipulable.

👉 Read Pentera’s analysis of AI alignment limits and value aggregation


Context

AI alignment is not only a model design problem. When an AI system must balance competing human preferences, fairness goals, safety expectations, and policy constraints, the issue becomes governance over value aggregation rather than just prediction quality or reasoning depth. That makes the article relevant to broader security and identity programmes because access policy, trust decisions, and control enforcement all depend on how conflicting requirements are resolved.

In identity-heavy environments, the same tension appears whenever human judgment, policy, and automated decisioning intersect. AI-assisted access reviews, fraud triage, and security prioritisation all rely on aggregation choices that cannot be made neutral by computation alone. Pentera’s conclusion is that this is a structural limit, not a tuning problem, and that is a familiar pattern in IAM and GRC programmes.


Key questions

Q: Why can’t AI alignment solve conflicting human values by optimisation alone?

A: Because the underlying problem is not just finding a better ranking rule. Arrow’s theorem shows that no aggregation method can satisfy every fairness and coherence requirement at once, and any real system must relax one of them. That makes value conflict a governance problem, not a pure optimisation problem.

Q: How should organisations handle strategic manipulation in human feedback systems?

A: Design the feedback process as if participants will learn how to game it, because the mathematics says they can. Reduce single-point influence, diversify reviewers, and validate high-impact outputs independently. The goal is not perfect honesty, but lowering the value of misreporting preferences.

Q: What do security teams get wrong about AI alignment?

A: Security teams often treat alignment as a one-time model training issue, then assume deployment controls will hold the line. That is wrong. Alignment is a runtime governance problem because prompts, data, tools, and feedback all change behaviour after release. The real control question is whether the organisation can observe, intervene, and roll back unsafe actions as they happen.

Q: Who should own AI value-setting decisions in an enterprise?

A: Ownership should sit with the business or security governance function that is accountable for the outcome, not with the model alone. If an AI system influences policy, access, or safety decisions, the organisation needs named human accountability for the objective function and its exceptions.


Technical breakdown

Arrow’s impossibility theorem and AI value aggregation

Arrow’s theorem says there is no preference aggregation rule that simultaneously satisfies universal domain, Pareto efficiency, independence of irrelevant alternatives, and non-dictatorship when there are at least three choices and two decision makers. The practical meaning is that a collective ranking cannot be both fully fair and fully coherent under all conditions. In AI alignment terms, the moment a system must combine multiple human value sets into one output, it inherits this structural conflict. Any design that appears to avoid the trade-off is silently relaxing one of the axioms, usually by narrowing the decision space or giving some preferences extra weight.

Practical implication: Practitioners should treat alignment schemes as explicit governance choices, not neutral optimisation layers.

Gibbard-Satterthwaite and strategic manipulation in preference elicitation

The Gibbard-Satterthwaite theorem shows that every non-dictatorial, onto social choice function with three or more alternatives is strategically manipulable. In plain terms, if the system is broad enough to choose among real options and not controlled by one person, someone can benefit by misreporting preferences. That matters for AI systems trained or tuned from human feedback, because the aggregator becomes part of the incentive structure. Once users, raters, or operators learn how the system interprets inputs, honest reporting is no longer guaranteed. The result is not merely bias. It is a predictable incentive to game the mechanism itself.

Practical implication: Design feedback pipelines assuming manipulation is inevitable, then reduce the value of gaming the input.

Why alignment-by-aggregation cannot be the final control

Alignment-by-aggregation assumes that if enough human input is collected, a model can converge on the right objective function. These theorems show why that assumption fails in the general case. Human values are inconsistent across contexts, and any mechanism that tries to reconcile them must choose between completeness, fairness, neutrality, and robustness against strategic behaviour. That is why RLHF, constitutional approaches, and voting-style preference systems can improve performance without eliminating the structural conflict. They are policy choices embedded in software, not final solutions to value conflict. In security terms, the objective function is itself part of the control surface.

Practical implication: Keep the objective function under governance review, especially where AI influences policy, access, or safety decisions.


Threat narrative

Attacker objective: The attacker’s objective is to steer the aggregated decision toward a preferred outcome by exploiting the mechanism’s strategic weakness.

  1. Entry occurs when an AI system is asked to collect and combine human preferences into a single decision rule.
  2. Escalation happens when the aggregation mechanism must choose between conflicting values, which forces the system to privilege some inputs over others.
  3. Impact is the emergence of a manipulable or inconsistent policy outcome that cannot be made universally fair and strategy-proof at the same time.

NHI Mgmt Group analysis

AI alignment exposes a governance ceiling, not a model-quality problem. Pentera’s argument matters because it shifts the question from whether a system can reason better to whether it can legitimately combine conflicting human values. Arrow’s theorem makes clear that no aggregation rule can preserve all desirable properties at once. For practitioners, the implication is that alignment is always a choice about which constraint to relax.

Value aggregation is the named concept here: the act of turning plural human preferences into one machine objective. That step cannot be made neutral by more compute, more data, or more feedback rounds. The Gibbard-Satterthwaite result adds the strategic dimension. Once participants know the mechanism, they can game it, which means governance must assume adversarial reporting, not ideal honesty. Practitioners should evaluate the mechanism as an incentive system, not just a learning loop.

AI systems that influence policy decisions inherit the same trade-offs as human decision processes. RLHF, constitutional AI, and voting-style protocols each encode a different compromise between fairness, stability, and manipulability. That makes the choice of alignment method an architectural governance decision, not a lab optimisation. Security and identity teams should insist on explicit ownership for these trade-offs before deployment.

The human role remains unavoidable because the choice of which axiom to violate sits outside the model. Pentera’s conclusion is strongest where it refuses the fantasy of a final social welfare function. Humans do not disappear when the model gets better. They move to the layer where priorities, exceptions, and acceptable trade-offs are set. Practitioners should frame AI oversight as policy authority, not post-hoc review.

What this signals

Value aggregation debt: enterprise AI programmes will keep accumulating hidden trade-offs unless they document which fairness, safety, or utility constraint is being relaxed. For identity and security teams, that means treating AI decision logic like any other policy control, with explicit ownership, review, and exception handling.

The practical signal is that high-impact AI workflows should be assessed for incentive sensitivity, not just accuracy. If a system’s outcome improves only when users behave perfectly, the control is brittle. That is the same governance lesson security teams already apply to privilege, approvals, and access certification.

This conversation also reinforces the need to keep AI governance connected to existing frameworks such as NIST AI RMF and NIST CSF, because the issue is accountability over decision making, not model scoring alone.


For practitioners

  • Define the alignment trade-off explicitly Document which property your AI decision process is allowed to relax: universality, neutrality, non-dictatorship, or strategy-proofness. That makes the governance boundary visible before deployment and avoids treating value aggregation as a hidden engineering choice.
  • Treat feedback channels as incentive surfaces Assume users, raters, or operators may adapt their inputs once they understand the system. Build controls that reduce the payoff from gaming, such as diverse review, randomisation, or independent validation of high-impact decisions.
  • Separate objective setting from model operation Keep the definition of fairness, safety, and acceptable override logic under human policy control rather than burying it inside model training. That separation is especially important where AI decisions affect access, risk scoring, or trust decisions.
  • Review AI governance through a control lens Map alignment mechanisms to existing governance structures such as approval authority, exception handling, and auditability. If the system cannot explain why one preference set overrode another, the programme has a governance gap, not just a modelling gap.

Key takeaways

  • AI alignment runs into a mathematical ceiling when it tries to aggregate conflicting human values into one objective.
  • Strategic manipulation is not an edge case in preference-based systems, it is a structural property that governance must anticipate.
  • Enterprises need explicit human ownership for objective setting, exception handling, and trade-off selection before AI makes policy-shaping decisions.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the technical controls, while ISO/IEC 27001:2022 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST AI RMFGOVERNThe article is about governance of AI value-setting and accountability.
NIST CSF 2.0GV.OV-01Governance and oversight are central because the topic is accountability for AI decisions.
NIST SP 800-53 Rev 5AC-1Policy and procedures are needed for AI decision authority and exception handling.
ISO/IEC 27001:2022A.5.15Access and decision authority need clear information security policy alignment.

Define AI decision policies and exception authority through documented control procedures.


Key terms

  • Arrow's Impossibility Theorem: A result in social choice theory showing that no rank-order voting system can convert individual preferences into a collective ordering while satisfying a full set of fairness conditions. It matters because any system that aggregates human values must relax at least one desirable property, even before implementation details are considered.
  • Gibbard-Satterthwaite Theorem: A theorem stating that any non-dictatorial, onto social choice function with three or more alternatives can be manipulated by participants who misreport preferences. In practice, it shows that collective decision systems cannot be both broadly expressive and fully strategy-proof.
  • Value Aggregation: The process of combining multiple human preferences, values, or constraints into one machine-readable objective or decision rule. In AI governance, this is where policy choices become operational, and where tensions between fairness, safety, and utility must be made explicit rather than hidden inside model behaviour.
  • Strategy-Proofness: A property of a decision mechanism where participants have no benefit from misrepresenting their preferences or inputs. It is a high bar that is often unattainable in collective AI and human feedback systems, which is why governance must assume some level of incentive distortion.

What's in the full article

Pentera's full article covers the mathematical detail this post intentionally leaves at the governance level:

  • Step-by-step walkthrough of Arrow’s theorem and the decisive coalition argument
  • The Gibbard-Satterthwaite manipulation theorem and why onto rules matter
  • How the article connects value aggregation limits to RLHF and constitutional AI
  • The closing interpretation of why humans remain the decision layer outside the model

👉 The full Pentera article walks through the theorems, the proof logic, and the implications for human decision making.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, IAM, and secrets management in environments where automated decisioning changes control boundaries. It helps security and identity practitioners connect governance design to the access and trust models their programmes depend on.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 17, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org