How Agencies Turn Responsible AI Policy Into Practice

A comparative governance guide to legal frameworks, case studies, and career pathways in public-sector AI oversight

By Max SheltonReviewed by PAP Editoral TeamUpdated September 13, 202625+ min read

What you’ll learn in this article…

  • EU hard law, US fragmented preemption, and Singapore soft law produce vastly different agency obligations.
  • Assign oversight ownership before deployment or no one signs the harm response.
  • Procurement contracts, not ethics statements, enforce responsible AI requirements on vendors.

Which government office is legally accountable when an algorithm denies someone's unemployment benefits or flags them for a tax audit? Most agencies cannot answer that question in writing, and that gap, not the underlying model, is what regulators and courts are now testing. A 2026 comparative study in the Cambridge Forum on AI: Law and Governance found that ethics statements rarely survive contact with procurement contracts, audit requirements, or due process claims.

This matters for public administration because the fix is organizational, not technical: lifecycle checkpoints, comparative governance models from the EU, US, and Singapore, and documented failures from the Netherlands to the IRS all point to the same design question. Who owns the decision when the system is wrong, and can that answer hold up under statutory scrutiny.

What 'Responsible AI' Means, and Who Actually Owns It

The difference between responsible AI that survives a court challenge and responsible AI that reads well in a press release comes down to one distinction: legal obligation versus ethical aspiration. Agencies that treat fairness, transparency, and accountability as optional design principles rather than enforceable requirements set themselves up for the kind of failures that make headlines.

From Ethics Statement to Statutory Duty

Principles like due process, reason-giving, and proportionality are not suggestions when they appear in administrative law. As Prema Elumalai and Ragul Olakkur Vijayakumar argue in their 2026 Cambridge Forum on AI: Law and Governance article Between Law and Code: Operationalizing Responsible AI in Public Administration, responsible AI must be built into law, procurement, deployment, and oversight, not layered on afterward as compliance theater. The practical implication: agencies need statutory power to audit algorithms, enforce compliance, and sanction misconduct. Contracts must require vendors to submit bias testing, impact assessments, and compliance results, effectively making constitutional and administrative norms into contractual obligations.

This framing organizes the rest of what agencies must do. Every AI system moves through a lifecycle: design, procurement, operational, and review. Responsible AI governance means attaching legal requirements and review checkpoints at each stage, not just at purchase.

The Ownership Problem No One Resolves

Who inside an agency actually owns AI oversight? The answer varies, and the variation itself creates risk.

  • CIO-led models concentrate technical expertise but often lack legal authority to enforce due process requirements or interpret administrative law constraints on automated decisions.
  • General-counsel-led models bring legal firepower but can bottleneck implementation and may lack fluency in how machine learning systems behave in production.
  • Program-manager-led models keep oversight close to mission delivery but distribute accountability so widely that no single office can be held responsible when something fails.

The Elumalai and Vijayakumar article describes a distributed governance model: statutory oversight powers for competent public bodies, procurement controls in contracting, and internal agency review mechanisms requiring human-signed decisions. Data protection authorities, privacy commissioners, and specialized bodies must be given statutory power to audit and enforce compliance.

Notice what this model does not do: it does not designate a single exclusive owner. Accountability requires feedback loops in both procurement and operational phases, meaning ownership is structural rather than positional.

Ambiguity as Governance Risk

For public managers and MPA students designing governance frameworks for AI in policymaking, the unresolved ownership question is not a neutral organizational choice. When no office clearly owns algorithm audits, bias testing, or grievance procedures, agencies default to the path of least resistance: treating responsible AI as someone else's problem until a failure forces clarity. The agencies that avoid scandals are the ones that assign accountability before they need to.

Three Governance Models: EU Hard Law, US Fragmented Preemption, and Singapore Soft Law

Not all responsible AI governance looks the same, and the model an agency operates under determines how much documentation, audit rigor, and legal exposure it must plan for. A 2026 comparative analysis published in the Cambridge Forum on AI: Law and Governance by Elumalai and Olakkur Vijayakumar systematically contrasts three jurisdictions that have taken fundamentally different paths: the European Union's unified hard-law regime, the United States' fragmented preemption landscape, and Singapore's facilitative soft-law approach. For public managers, the practical question is not which model is philosophically superior but which obligations actually bind their agency and how deep the compliance architecture needs to go.

DimensionEuropean Union (AI Act)United States (Fragmented Preemption)Singapore (Model AI Governance Frameworks)
Legal basisDirectly regulated under the EU Artificial Intelligence Act. High-risk AI systems listed in Annex III face binding obligations, with timelines partly deferred by the Digital Omnibus amendments (Regulation (EU) 2026/1744). AI embedded in regulated products (Annex I) carries additional product-safety requirements.No single comprehensive AI statute. Governance relies on a patchwork of federal executive guidance, sector-specific statutes (consumer protection, civil rights, financial regulation), and state-level laws. Federal rules can preempt or constrain stricter state approaches.Non-binding frameworks issued by the Personal Data Protection Commission and IMDA, covering general AI, generative AI, and agentic AI. These complement statutory regimes such as the Personal Data Protection Act but do not themselves create enforceable AI-specific legal duties.
Primary enforcement mechanismMandatory conformity assessment, technical documentation, CE marking, and registration in the EU AI database. Standalone high-risk systems must meet these requirements by 2 December 2027; AI in regulated products by 2 August 2028. Non-compliance triggers supervisory enforcement and sanctions.Existing regulators (FTC, CFPB, EEOC, state attorneys general) apply general consumer-protection, civil-rights, and financial laws to AI. Emerging federal and state rules create overlapping obligations, but no unified AI enforcement body exists.Compliance is voluntary. Regulators reference the frameworks when setting sectoral expectations, and organisations are encouraged to publish documentation on training data, evaluation results, risk mitigations, and model limitations. No statutory fines attach to the AI governance frameworks themselves.
Public-sector obligationsAgencies using high-risk AI in areas such as employment, biometrics, migration, or law enforcement must ensure no system is placed on the market or put into service without prior conformity assessment. Legacy systems deployed before August 2026 receive an extended transition through 2 August 2030.Government AI use is governed by a mix of federal agency policies and state-level procurement and transparency rules. No single nationwide statutory framework applies specifically to government AI systems, leaving agencies to assemble compliance from multiple overlapping authorities.Government agencies are expected to follow framework principles: data governance, transparency, human-in-the-loop oversight, risk assessment, and documentation of model operations. For agentic AI, agencies should assess and bound risks upfront. All expectations are framed as policy guidance, not enforceable legal duties.
Implied documentation and audit rigorHighest. Agencies must maintain full technical documentation, conduct conformity assessments, and register each high-risk system in a public database. The binding deadline structure creates a clear minimum compliance floor with legal consequences for gaps.Variable and often reactive. Without a unified standard, agencies face legal risk from multiple enforcement vectors but lack a single checklist. Documentation rigor depends on the specific federal or state rules that apply to each use case, creating inconsistency across jurisdictions.Moderate but discretionary. Agencies are encouraged to document risk assessments, model limitations, bias mitigations, and user-data protections. Because adoption is voluntary, the depth of an agency's audit trail reflects internal policy commitment rather than statutory mandate.
Key practical takeaway for public managersBuild compliance into procurement from day one. Vendors must deliver conformity-ready systems, and agencies bear deployer obligations. Timelines are fixed and enforceable.Conduct a cross-jurisdictional legal scan before deploying any AI system. The absence of a unified framework increases, rather than reduces, legal risk because multiple regulators may assert authority simultaneously.Treat the frameworks as a governance floor, not a ceiling. Voluntary adoption today can become a baseline expectation tomorrow, and documented compliance positions an agency well if statutory obligations emerge.

Mapping Administrative Law to the AI Lifecycle

Administrative law rarely maps to one moment in an automated decision. The comparative framework organizes due process, transparency, reason-giving, and proportionality across four phases: design, procurement, operational, and review.1 That lifecycle view matters for public policy making because it tells agencies what to require before a vendor is selected, not only after a model is live.

Due process moves from notify, defer, and human-readable reasons in design to logging, human review, and explainability in procurement. Once a system is operational, due process requires notice, contest opportunities, and reasoned decisions; ex post review adds traceability through logs and model versions.1 The newer due-process standards push the same direction: traceability and human oversight.2

Transparency requirements phase similarly: documentation and algorithmic transparency at design; model cards, impact assessments, and registries at procurement; front-end transparency and contextual tailoring in operation; and tiered disclosure ex post.1 Contextual tailoring, in particular, varies transparency by context and risk.

Reason-giving maps to explanation techniques in design, explanation artifacts and official training in procurement, a “why this tool, use, outcome” requirement in operation, and rationality review ex post.1 Proportionality imposes risk-sensitive design, fundamental-rights impact assessments, tiered procedural protections, and proportionality review.1 For public administration education, the practical lesson is that these are not abstract ethics: they are enforceable constraints.

The techno-legal architecture synthesized from UNESCO, the EU AI Act, and the AI Bill of Rights makes the mapping concrete: algorithmic registries and audit trails; procedural entitlements such as notice, hearing, access, reasons, and review; front-facing, evidentiary, and technical transparency layers; deployment controls such as fundamental-rights impact assessments and proportionality; and reflexive governance for ongoing recalibration.1 A fundamental-rights impact assessment runs through planning, risk assessment, and monitoring.4 Courts are adding their own disclosure expectations, requiring agencies to state AI use, system, version, limitations, stage, and purpose.5

Taken together, these legal mappings produce accountability through the lifecycle rather than at a single approval point.6 For public managers, the payoff is not a policy statement but an evidence-based decision making in government framework for auditing vendors and internal systems before decisions cause harm. That turns responsible-AI commitments into pre-deployment checks, not post-harm apologies.

Algorithmic Impact Assessments and Audit Trails: The Core Checklist

When practitioners search for concrete guidance on how agencies audit AI models after deployment, the algorithmic impact assessment is the single most actionable mechanism available. Federal memoranda, particularly OMB M-24-10 and the later M-25-21, spell out what agencies must document, when, and for whom. The checklist below distills those obligations into the core components every public manager, policy analyst, or MPA student should know.

  • Use-Case Description and Intended Purpose
    Every covered AI system must be documented with a clear statement of its intended purpose and expected benefit, supported by metrics or qualitative analysis. Under M-25-21, this documentation is required before deployment for any high-impact AI use case and must be updated periodically throughout the system's lifecycle.
  • Risk Tier Classification
    Agencies must determine whether each AI use case is rights-impacting or safety-impacting. M-24-10 requires agencies to document and recertify these determinations, and the resulting classification drives which minimum risk-management practices apply. This tiering is the gateway to every downstream obligation.
  • Bias Testing and Data Quality Documentation
    Impact assessments must document data used in design, development, training, testing, and operation, including how data was collected and prepared. M-25-21 also requires agencies to document potential impacts on privacy, civil rights, and civil liberties, and to note whether data will be publicly disclosed as an open government data asset.
  • Human-Review Pathway
    For rights-impacting and safety-impacting AI, agencies must establish documented human oversight controls. This means specifying who reviews automated determinations, under what conditions a human intervenes, and how escalation procedures are recorded, ensuring that no consequential decision relies entirely on an opaque model.
  • Retirement and Sunset Triggers
    A responsible AI lifecycle does not end at deployment. Agencies should define clear conditions under which a system is decommissioned, whether due to performance drift, legal changes, or failure to pass periodic reassessment. M-25-21's requirement for ongoing lifecycle updates creates a natural checkpoint for these decisions.
  • Audit-Trail Records and Retention
    Federal agencies must maintain audit-trail records showing the basis for determinations, testing results, monitoring outputs, and any lifecycle changes. These records support both internal accountability and external review. M-24-10 requires agencies to publish annual AI use-case inventories and provide public notice with plain-language documentation of covered uses.
  • Inventory Reporting and Public Transparency
    Under M-24-10, agencies submit AI use-case inventories to OMB annually and publish them publicly, with limited exceptions. This transparency requirement means that the public, and oversight bodies, can see which AI systems an agency operates and whether impact assessments are in place. The initial inventory submission deadline was December 2024, establishing the federal baseline.

Public Sector AI Procurement Requirements: Translating Rules Into Vendor Contracts

Responsible AI principles only matter if they survive contact with a purchase order. The table below maps governance requirements to concrete contract language that agencies can adapt for RFPs, task orders, and blanket purchase agreements. These obligations draw on federal guidance from OMB and GSA, as well as published state and local procurement frameworks from Connecticut and the District of Columbia. When a vendor system fails a required risk or bias review after award, the contractual remedy depends on how clearly these clauses were written: agencies with explicit right-to-audit and performance-monitoring provisions can compel corrective action, suspend use, or terminate for cause, while vague "best efforts" language leaves the agency exposed.

Requirement CategoryWhat It Means for AgenciesVendor or Contract Obligation
Risk Management and Bias TestingAgencies must evaluate AI buys against explicit risk, data, privacy, and performance controls, per OMB M-25-22. High-impact use cases require minimum risk management practices built into the acquisition.Contract terms must require minimum risk management practices for high-impact AI, including ongoing testing and monitoring of performance, risk, and effectiveness throughout the period of performance.
Data Provenance, Privacy, and Intellectual PropertyAgencies should secure operational independence and reuse rights so AI systems remain controllable after award. The U.S. Department of the Interior requires sufficient rights in custom-developed code and data stored on acquired systems.Contracts must address intellectual property rights, government data use, privacy protections, and vendor lock-in safeguards such as knowledge transfers, data and model portability, and sufficient IP and data rights.
Model Documentation and Decision TransparencyConnecticut requires agencies to verify a vendor's decision-making transparency before purchase. This creates a procurement gate: no disclosure, no contract.Vendors must disclose required information about how AI models reach decisions. Agencies may not procure a solution without verifying that the vendor has met disclosure requirements and that decision-making transparency has been confirmed.
Pilot Testing and Technical ValidationGSA guidance directs agencies to use testbeds, sandboxes, or pilot programs before committing to large-scale purchases. The AI Center of Excellence recommends Statements of Objectives when the solution path is uncertain.Solicitations should include technical tests as evaluation criteria. Vendors must participate in sandbox or pilot validation, and agencies should structure buys around measurable outcomes rather than overly prescriptive specifications.
Ongoing Monitoring and Vendor NotificationOMB M-25-22 requires that agencies monitor vendor performance continuously, not just at the point of award. Vendors must alert the agency to changes in AI capabilities after delivery.Contracts must include vendor performance requirements, consumption reviews, and a notification obligation whenever the vendor introduces new AI enhancements, features, or components that alter system behavior.
AI Safeguarding Clause InsertionGSA's Federal Acquisition Service has proposed a Basic Safeguarding of Artificial Intelligence Systems clause for use across solicitations. Agencies adopting AI-specific terms must incorporate explicit safeguarding controls.The Basic Safeguarding of Artificial Intelligence Systems clause must be inserted into solicitations and contracts for AI capabilities, ensuring baseline protections are standardized across procurement instruments.
Cross-Method Procurement CoverageThe District of Columbia's AI Procurement Handbook requires AI governance obligations in solicitations for services and information technology regardless of procurement method, preventing gaps when agencies use simplified or micro-purchase channels.AI procurement requirements apply across all solicitation types and buying channels. Agencies must normalize AI controls so that governance is not limited to a single procurement method.
Interagency Collaboration and Use-Case ScopingState procurement collaborative guidance (NASPO and NASCIO) urges agencies to start with targeted use cases and foster collaboration between procurement and IT leadership, including CIO, CAIO, CDO, CISO, and privacy officials.Vendors should expect evaluation criteria tied to specific mission needs rather than generic AI capability claims. Agencies convert governance into vendor requirements by defining policies, narrowing use cases, and aligning procurement with IT oversight.

When Responsible AI Fails, and When It Works: Dutch Benefits, UK Exams, IRS Audits, and Estonia's Kratt

Abstract principles like transparency and proportionality, central to any public policy definition, mean little until you see what happens when they are absent. Four cases, three failures and one working model, show public managers exactly where the design choices go wrong and what building in an appeal pathway from the start actually looks like.

The Dutch Childcare Benefits Scandal

From roughly 2004 to 2019, the Dutch tax authority used a risk-scoring system to flag childcare-benefit claims for fraud enforcement. Amnesty International later concluded that racial and ethnic profiling was baked into the model's design1, and low-income parents and caregivers, disproportionately from ethnic minority backgrounds, were falsely accused of fraud and ordered to repay benefits, often tens of thousands of euros, with no meaningful path to contest the algorithm's judgment. Estimates of the scale vary: one accounting puts wrongly accused families at 26,000 as of 20132, while a broader comparative review cites 35,000 welfare recipients whose fundamental rights were infringed3. The consequences were severe: families lost homes, went into crushing debt, and more than 2,000 children were placed in foster care amid the resulting household instability. The government later promised affected families 30,000 euros in compensation4. The December 2020 Ongekend Onrecht inquiry found that the system violated basic rule-of-law principles, and the Rutte III cabinet resigned in January 20214.

UK A-Level Grading and the Standardization Backlash

When Covid canceled 2020 exams, England's Ofqual built a standardization algorithm to moderate teacher-predicted grades using historical school performance data. Announced August 13, 2020, the results downgraded nearly 36 percent of predictions by one grade and 3 percent by two grades5, roughly 280,000 entries adjusted down by at least one grade in England alone6. Only about 59 percent of grades matched teacher predictions, and about 2 percent moved up7. Because the model leaned on school-level history to make school quality assessments, it systematically penalized strong students at historically lower-performing schools, a fairness problem that triggered immediate public backlash. Ofqual reversed course on August 17, nullifying the algorithmic grades in favor of unmoderated teacher assessments8, which pushed top grades up sharply (Ofqual estimated an unprecedented 12.5 percentage point jump in A and A* awards without moderation)6.

IRS Audit Selection and the Disparate Impact Question

The US offers a less resolved version of the same problem. Automated case-selection tools used to flag tax returns for audit have drawn sustained scrutiny over whether they produce disparate outcomes across income and demographic lines, with researchers and lawmakers pressing the IRS for more transparency about how selection models work and who gets audited as a result. Unlike the Dutch and UK cases, there has been no single dramatic reversal or resignation, partly because the audit-selection process itself remains far less visible to the public. That opacity is itself the finding: without a published methodology or a clear appeal mechanism, disparate-impact concerns linger unresolved rather than getting adjudicated and closed.

Estonia's Kratt: Building the Appeal In From the Start

Estonia's Kratt initiative, a national framework for interoperable AI-powered virtual assistants across government services, points the other direction. Rather than deploying a single opaque model, Estonia built its digital-government AI on a shared legal and technical architecture designed for interoperability across agencies, with accountability and rights protections treated as part of the system's architecture rather than bolted on afterward.

The Common Thread

For students pursuing an MPA or MPP Degree, every failure here shares two defects: an opaque model whose logic wasn't disclosed to the people it affected, and no working appeal pathway to challenge an automated decision before real harm occurred. Estonia's approach inverts both defects by design, not by accident.

State and Municipal AI Governance: Where the Action Really Is

California's Executive Order N-5-26 does more work in a single directive than most federal guidance has managed in years: it instructs the California Department of Technology, the Department of General Services, and other state agencies to build procurement certifications, vendor safeguards, and formal AI governance processes for state use of the technology. Agencies were told to act immediately, and later state announcements described follow-on tools and programs rolling out under that mandate. For MPA and MPP students, this is the clearest current example of public policy for AI built directly into procurement rather than bolted on afterward.

New York City's Oversight Office Model

The New York City Council approved local legislation creating an Office of Algorithmic Accountability, paired with compliance standards that city agencies must meet when deploying AI systems.1 Unlike a voluntary task force, this is a council-approved local law with immediate effect, meaning the oversight office isn't an advisory afterthought but a standing compliance function inside city government. That distinction matters for anyone assessing which governance structures actually survive contact with budget cycles and agency turnover: an office with statutory footing tends to outlast the political moment that created it.

Seattle and San Jose: Governance in Motion, Status Unclear

Not every city has reached the same clarity. Seattle has pursued municipal AI oversight and algorithmic accountability efforts, but the operational status of that oversight remains unclear as of current reporting. San Jose has likewise moved toward board-style AI governance, yet no clear current operational status has emerged from available public records. These two cases are instructive precisely because they are unresolved: they show that announcing a governance mechanism and operationalizing one are different milestones, and a program manager evaluating a peer city's model through program evaluation in public administration should ask which stage that city has actually reached.

The Common Thread

Across all four jurisdictions, the mechanism, whether executive order, oversight office, registry, or board, is only as strong as its enforcement pathway. California ties its mandate to procurement leverage. New York City ties its office to statutory compliance standards. Seattle and San Jose illustrate what happens when the mechanism exists on paper before the enforcement infrastructure catches up. For students and practitioners building AI governance in their own jurisdictions, the practical lessons in implementation sequencing sit at the state and municipal level, not in federal memoranda.

Measuring Responsible AI Maturity: The Stages Agencies Move Through

Agencies at every level of government ask the same question: how do we know our responsible AI efforts are actually working? A maturity model provides a concrete answer. Rather than treating compliance as a binary pass/fail, the framework below maps four sequential stages, each with a specific marker that proves an agency has moved beyond good intentions into demonstrable governance.

Four-stage responsible AI maturity model for government agencies, from ad hoc pilot use through continuous KPI monitoring

Building Responsible AI Capacity: Workforce, Training, and Careers in Public Administration

Agency capacity for responsible AI is not about hiring a room of data scientists. It is the combination of skills, roles, and routines that lets a public organization buy, build, audit, and explain automated systems without losing control of due process, transparency, or fairness.

The Skill Set Is More Translator Than Coder

Agencies increasingly need staff who can read a vendor's model card, run a basic bias test on a sample output, and translate a legal standard such as "reasoned decision-making" into a concrete procurement requirement. Most of this work is not building algorithms. It is interpretation: reading the technical artifact well enough to know which questions to ask, then writing contract language a vendor cannot quietly walk around. A policy analyst who can spot a missing data-quality check is often more valuable than a junior data scientist. Public administration programs have begun adding AI governance modules to MPA and MPP curricula, often inside technology policy, administrative law, or program evaluation courses. Students should look for coursework that pairs case studies such as the UK exam-grading and Dutch childcare-benefits failures with hands-on exercises like drafting an algorithmic impact assessment or reviewing a model card.

Cross-Training Beats New Headcount

Budget realities mean few agencies will create a new AI oversight office from scratch. The more common path is cross-training existing analysts, procurement officers, program managers, and legal staff. A budget analyst who learns to read a fairness report, or a contract specialist who can specify audit-log retention requirements, creates capacity without adding net-new positions. Some agencies rotate staff through a short AI governance fellowship or embed a data steward in the chief information officer's shop for a quarter. This approach is slower than hiring, but it builds institutional memory that outsourced consultants rarely leave behind.

Accountability Must Exist Before Training Pays Off

Capacity-building only compounds when someone is formally responsible for it. The accountability gap described earlier, between the CIO, general counsel, and program managers, can turn every training session into scattered awareness with no durable practice. The clearest agencies name a responsible AI lead or designate a unit to own the checklist, and they write that ownership into job descriptions and performance reviews before spending on new tools or courses. Without that named owner, a model card is just another unread attachment.

Recent News

Recent Articles