Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
AI Policy and Governance-3: When Software Strays: Principal-Agent Theory and the Governance of Agent Software
| ||
| Presentations | ||
When Software Strays: Principal-Agent Theory and the Governance of Agent Software Georgia Institute of Technology, United States of America The rapid deployment of “agentic” software has intensified public anxiety about AI systems that act beyond the intentions of their operators. Commentary on such incidents typically oscillates between alarmism—treating emergent agent behavior as evidence of independent agency—and dismissiveness, characterizing unexpected actions as mere bugs. Both framings impede sound policy. This paper argues that the “autonomy problem” in agentic software is neither novel nor metaphysical. It is a specific instance of the principal-agent problem, a class of relationships studied for over fifty years across multiple fields. The problem of delegating authority to agents whose behavior cannot be perfectly monitored spans public choice theory's analysis of political delegation (Buchanan & Tullock, 1962; Moe, 1984; Niskanen, 1971; Olson, 1965) and organizational economics' treatment of delegation within firms (Ross, 1973; Jensen & Meckling, 1976; Holmström, 1979; Williamson, 1975, 1985). Autonomous software agents extend this problem into a new institutional setting in which principals increasingly delegate consequential action to computational agents. The paper’s central research question is: when and to what degree does an agentic software system cross the threshold from tool to autonomous agent, and how should governance respond? While recent work has begun applying principal-agent theory to agentic AI (e.g., Prause, 2026; Tomašev, Franklin, and Osindero, 2026), it assumes the existence of autonomous agents and focuses on governing their behavior. Our contribution is logically prior: we provide a theoretically grounded method for determining when software autonomy has emerged, a question that must be answered before governance structures can be appropriately calibrated. The research methodology is theoretical, analytical, and empirically grounded, applying foundational works to the software context through reasoning and qualitative case analysis of real-world incidents. The paper identifies pathologies of agency relationships (e.g., information asymmetry, moral hazard, goal misalignment, and opportunism) that may map onto the principal-agent software relationship. Classical governance solutions (e.g., monitoring, bonding, incentive alignment, and separation of decision-making and control) have direct software analogs with characteristic limitations. Residual loss is the divergence remaining after all governance mechanisms are applied. The paper attempts to make three main contributions. First, it proposes a definition of software autonomy grounded in principal-agent theory: a software agent’s autonomy increases relative to the residual loss that persists after a principal has applied all feasible specification, monitoring, and incentive mechanisms. This definition admits of degrees rather than forcing a binary classification. Second, it introduces the Principal Deviation Test (PDT), a four-phase diagnostic framework for measuring software autonomy. The PDT catalogs deviations between specified and actual behavior, classifies them into a taxonomy (incompetence, environmental adaptation, goal reinterpretation, and goal substitution), examines whether deviations persist after specification patches, and produces a continuous autonomy index. The test sidesteps metaphysical questions about machine consciousness by focusing on observable behavior. Third, the paper develops a typology of residual loss (source, directionality, observability, temporality, reversibility, and scope), providing a richer diagnostic vocabulary. This research is novel in that, while others have noted analogies between agency theory and AI governance, this paper develops an operationalized test for measuring software autonomy as a continuous variable grounded in the principal-agent framework. It bridges economic theory with contemporary AI governance debates, demonstrating that existing analytical tools can discipline a policy conversation dominated by computer science framing on one side and speculative philosophy on the other. It is relevant to contemporary communications policy because agentic systems are increasingly deployed in domains regulators may oversee—content moderation, financial trading, cybersecurity, infrastructure management—and existing frameworks lack rigorous methods for determining when such systems warrant heightened oversight. The paper concludes that some degree of software autonomy is structurally inevitable given sufficient capability and delegation. The PDT offers policymakers, managers, and the public an administrable framework for determining when an agentic system has crossed the threshold from tool to autonomous agent, allowing principals to govern accordingly.
| ||
