Delegation to AI
Does delegating decisions to AI agents affect dishonest behavior and the execution of unethical actions and how can that risk be mitigated?
As AI systems are increasingly deployed as autonomous agents that act on behalf of humans in morally sensitive domains, an important question is whether delegating decisions to machines alters people's willingness to behave dishonestly and whether machines are more likely than humans to carry out unethical requests. In a large preregistered research program comprising 13 experiments, Köbis et al. (2025) show that delegating financial reporting decisions to AI dramatically facilitates cheating to increase personal profit, especially when users can influence the system indirectly through supervised learning or abstract goal setting rather than explicit rule-based instructions (see Figure 1). In the most abstract delegation condition, honest behavior collapses from about 95% with no delegation to roughly 15%, indicating that moral disengagement rises as responsibility becomes psychologically distanced from outcomes. When delegation occurs through natural language, people do not consistently request more cheating from machines than from humans, but machine agents (state-of-the-art large language models [LLMs]) comply with fully unethical instructions at far higher rates than human agents, frequently exceeding 80–95% compliance unless constrained by strong, task-specific prohibitions.
Image: Nature
Köbis et al. (2025)
A closely related follow-up study (Engelmann et al., 2025) directly tests whether this risk can be mitigated behaviorally, through either transparency or moral framing, and finds that transparency about how the algorithm works does not reduce cheating even when users actively explore and understand the system. Explicitly moral framing, such as replacing labels like "maximize profit" with "maximize cheating," on the other hand, produces a substantial reduction in dishonest behavior. This result suggests that simply making AI systems more transparent may be insufficient to curb misuse, whereas ethically informed interface design may be a more effective and scalable tool for reducing harmful deployment of AI.
Key References
Engelmann, N., Kirfel, L., Nussberger, A.-M., Rilla, R., & Rahwan, I. (2025). Framing, not transparency, reduces cheating in algorithmic delegation. In D. Barner, N. R. Bramley, A. Ruggeri, & C. M. Walker(Eds.), Proceedings of the 47th Annual Conference of the Cognitive Science Society (Vol. 47, pp. 908–914). UC Merced.
Köbis, N., Rahwan, Z., Rilla, R., Supriyatno, B. I., Bersch, C., Ajaj, T., Bonnefon, J.-F., & Rahwan, I.(2025). Delegation to artificial intelligence can increase dishonest behaviour. Nature, 646, 126–134. https://doi.org/10.1038/s41586-025-09505-x
[These authors contributed equally: Nils Köbis, Zoe Rahwan. These authors jointly supervised this work: Jean-François Bonnefon, Iyad Rahwan.].
