Join our Newsletter — 33% off our NHI Course
Home FAQ AI Security Why can’t AI alignment solve conflicting human values…
AI Security

Why can’t AI alignment solve conflicting human values by optimisation alone?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated August 17, 2026 Domain: AI Security

Because the underlying problem is not just finding a better ranking rule. Arrow’s theorem shows that no aggregation method can satisfy every fairness and coherence requirement at once, and any real system must relax one of them. That makes value conflict a governance problem, not a pure optimisation problem.

Why This Matters for Security Teams

ai alignment often gets described as a tuning problem, but conflicting human values are rarely resolvable by a single objective function. In practice, teams are trying to balance safety, utility, fairness, privacy, and accountability under changing conditions. That is closer to governance than optimisation, because every choice about what the system should prioritise creates a tradeoff somewhere else. The issue is especially important in high-stakes environments where model behaviour affects access, eligibility, triage, or enforcement decisions.

Security and risk leaders should treat alignment as a control problem with policy inputs, review gates, and escalation paths, not just a model-training outcome. Current guidance across frameworks such as the NIST Cybersecurity Framework 2.0 reinforces that effective governance depends on defined roles, ongoing oversight, and measurement, not a one-time technical fix. The same logic applies to AI systems that make or influence decisions on behalf of an organisation. If stakeholders disagree on what counts as harm, fairness, or acceptable error, optimisation can only hide the conflict inside the loss function.

In practice, many security teams encounter alignment failures only after an AI system has already made a contested decision at scale, rather than through intentional governance design.

How It Works in Practice

Real alignment work starts by making value tradeoffs explicit. That means separating the technical question of model performance from the policy question of which outcomes are acceptable, who decides, and what happens when values conflict. The best practice is evolving, but most mature programs now use a combination of value statements, risk acceptance criteria, human oversight, and post-deployment monitoring. A model can optimise for one metric, such as accuracy or helpfulness, while still violating organisational expectations if those expectations were never translated into concrete constraints.

For AI systems used in regulated or sensitive contexts, this often requires an operating model with several layers:

  • Define the decision domain and the harms that matter most, such as false positives, discriminatory outcomes, unsafe content, or overconfident automation.
  • Convert policy intent into measurable controls, review thresholds, and escalation triggers.
  • Validate outputs against test cases that reflect competing stakeholder views, not just benchmark performance.
  • Monitor for drift, because value assumptions can change faster than model weights.

Where AI is used in security operations, governance should also consider how model recommendations interact with human approval, incident response, and identity controls. A model that advises access changes, blocks transactions, or supports triage needs clear accountability boundaries and auditability. NIST AI risk guidance and model governance practices increasingly stress that alignment requires lifecycle management, including design, testing, deployment, and continuous review. Research-led standards such as NIST AI Risk Management Framework and adversarial AI threat work from MITRE ATLAS both point to the same operational truth: optimisation does not remove disagreement, it only formalises one interpretation of it.

These controls tend to break down when organisations deploy general-purpose models into ambiguous decision processes because no one has authority to resolve tradeoffs when the model’s objective conflicts with business, legal, or ethical requirements.

Common Variations and Edge Cases

Tighter alignment constraints often increase model complexity, oversight cost, and latency, requiring organisations to balance consistency against operational speed. That tradeoff becomes sharper when the system must serve multiple jurisdictions, customer groups, or policy regimes. There is no universal standard for how to resolve incompatible values in every context, so the governance model matters as much as the training method.

One common edge case is when teams assume a single ranked preference list can represent all stakeholders. That works poorly when values are in genuine conflict, such as safety versus creativity, privacy versus observability, or consistency versus flexibility. Another edge case appears in agentic systems, where the model can plan, act, and call tools. In that setting, the problem is not just what the model predicts, but what authority it has to act on those predictions. This is where alignment overlaps with identity and access governance, especially for approvals, tool permissions, and audit trails.

Frameworks such as NIST AI Risk Management Framework and OWASP guidance for AI and agentic systems support the view that systems should be constrained, monitored, and challenged continuously. The practical takeaway is simple: optimisation can help choose among known tradeoffs, but it cannot eliminate disagreement about which tradeoffs are acceptable. When that disagreement is hidden, the model appears aligned until a real-world case exposes the unresolved value conflict.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATLAS and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST AI 600-1 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFAI governance and lifecycle risk management fit this value-conflict problem.
NIST CSF 2.0GV.OCOrganisational context and governance are central when values conflict.
MITRE ATLASAdversarial manipulation and model misuse can worsen unresolved alignment issues.
OWASP Agentic AI Top 10Agentic systems need guardrails when autonomous actions depend on contested objectives.
NIST AI 600-1GenAI profile guidance helps operationalise safety, transparency, and oversight.

Restrict tool use, require approvals, and monitor actions that could amplify value conflicts.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on August 17, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org