The Governance Gap in Educational AI
Artificial intelligence is entering classrooms at a pace that governance frameworks were not designed to match. Adaptive learning platforms, AI-powered tutoring systems, automated essay graders, behavioural analytics tools, and generative AI assistants are being deployed across schools and universities worldwide — often with minimal scrutiny of how they handle the sensitive personal data of students, many of whom are minors.
The result is a governance gap of significant proportions. The primary legal framework governing student data in the United States — the Family Educational Rights and Privacy Act (FERPA) — was enacted in 1974, before the internet existed, let alone machine learning. In the European Union, the General Data Protection Regulation (GDPR) provides stronger protections but was not designed with AI-specific risks in mind. Neither framework adequately addresses the distinctive challenges posed by AI systems that can infer psychological states from typing patterns, build behavioural profiles from interaction data, and use student inputs to train commercial models.
This data brief maps the current landscape: the scale of AI deployment in education, the specific governance failures that have emerged, the fragmented policy responses taking shape, and what genuine data sovereignty in education would require.
Key Data Points: The Scale of AI in Education
Deployment and Policy Fragmentation
- As of mid-2026, at least 35 US states and Puerto Rico have issued official AI guidance for K-12 schools — but guidance is not regulation, and compliance is not mandatory.
- Major districts including New York City Public Schools and Los Angeles Unified have implemented significant restrictions or bans on student-facing generative AI, while others such as Katy ISD have adopted "scaffolded" access models with progressive, supervised deployment.
- The US Department of Education has signalled a shift towards outcomes-based contracting, evaluating AI tools based on measurable student learning outcomes rather than usage metrics — but this framework remains aspirational rather than operational.
- Research collaborations between Carnegie Mellon University, OpenAI, Stanford, and the University of Tartu are working to establish benchmarks for AI in education, but peer-reviewed evidence of efficacy remains limited.
The Shadow AI Problem
- A significant proportion of AI usage in schools occurs via "shadow AI" — teachers and students using unauthorised consumer-grade tools outside district oversight and data agreements.
FERPA was enacted in 1974 — before the internet, before smartphones, before machine learning. Applying it to AI systems that can infer psychological states from typing patterns is not a compliance exercise; it is an act of institutional imagination that the law was never designed to support.
- Shadow AI systematically bypasses district-level data processing agreements, leading to potential FERPA violations through the unauthorised disclosure of personally identifiable information (PII).
- Technology providers are actively embedding themselves into district infrastructure through free, enterprise-level access to AI tools, often with FERPA-aligned data agreements — but with concerns about vendor lock-in once promotional access periods conclude.
The FERPA Architecture and Its AI-Era Failures
FERPA grants students (and parents of minor students) the right to access, correct, and control the disclosure of their education records. It prohibits educational institutions from disclosing personally identifiable information from education records without consent, with exceptions for "school officials" who have a "legitimate educational interest" in the information.
FERPA was enacted in 1974 — before the internet, before smartphones, before machine learning. Applying it to AI systems that can infer psychological states from typing patterns is not a compliance exercise; it is an act of institutional imagination that the law was never designed to support.
The "school official" exception has become the primary mechanism through which educational institutions attempt to authorise AI vendor access to student data. To qualify, an AI vendor must be under the direct control of the institution, perform a function the school would otherwise handle with employees, and be limited to a "legitimate educational interest." In practice, applying this exception to AI systems raises questions that FERPA's drafters could not have anticipated.
Model Training: The Critical Vulnerability
The most significant AI-specific risk under FERPA is model training. When students interact with an AI tutoring system, their inputs — essays, questions, responses, behavioural patterns — may be used by the vendor to train or improve their AI models. This generally violates FERPA's purpose-limitation principle: student data disclosed for an educational service cannot be repurposed for commercial model development without explicit consent.
The challenge is that model training is often opaque. Data processing agreements may prohibit explicit model training while permitting "service improvement" activities that are functionally equivalent. Institutions lack the technical expertise to audit vendor practices, and the regulatory enforcement mechanisms available under FERPA — which can ultimately result in the withdrawal of federal funding — are too blunt to address nuanced compliance failures.
Algorithmic Profiling and AI-Generated Education Records
A second critical vulnerability concerns algorithmic profiling. AI systems deployed in educational settings generate inferences about students — assessments of learning styles, behavioural patterns, emotional states, and academic potential — that may themselves constitute education records under FERPA. If these AI-generated inferences are used to make consequential decisions about students (placement, discipline, special education services), they must be subject to the same access, correction, and disclosure rights as traditional education records.
Emerging research highlights the risks of algorithmic bias in these systems: AI models trained on historical educational data may encode and amplify existing inequities, producing discriminatory outcomes for students from marginalised backgrounds. The governance frameworks needed to audit, challenge, and correct these algorithmic decisions do not yet exist in most educational institutions.
The Sovereignty Deficit: Beyond Data Residency
Shadow AI — teachers and students using unauthorised consumer-grade tools outside district oversight — is not a marginal phenomenon. It is the dominant mode of AI adoption in education, and it is systematically bypassing every data protection framework institutions have built.
The concept of "data sovereignty" in education has historically been interpreted narrowly — as a question of where data is stored and processed. Institutions have sought to ensure that student data remains within national or regional boundaries, subject to local law, rather than being processed in foreign jurisdictions under foreign legal frameworks.
This narrow interpretation is insufficient for the AI era. The 2026 Microsoft Digital Sovereignty Summit for education leaders identified a more demanding conception of sovereignty: one that requires transparency not merely about data location, but about how AI models are trained, where prompts are processed, who maintains access to the data lifecycle, and how AI-generated inferences are used in decision-making.
The shift from data residency to genuine data sovereignty in education requires transparency about how AI models are trained, where prompts are processed, and who maintains access to the data lifecycle — questions that most EdTech contracts do not answer.
This expanded conception of sovereignty has practical implications for institutional procurement and governance. A "sovereign cloud" solution that stores data within national boundaries but allows the vendor to use interaction data for model training does not provide genuine data sovereignty. Genuine sovereignty requires contractual and technical controls over the entire data lifecycle — from collection through processing, inference, storage, and deletion.
The Workload-Specific Approach
Leading institutions are moving away from one-size-fits-all data governance architectures towards workload-specific approaches that apply tailored controls based on the sensitivity of the data and the risk profile of the AI application. A student information system containing grades, attendance records, and IEP data requires different governance controls than an AI writing assistant used for low-stakes classroom exercises.
This workload-specific approach requires institutions to develop the technical and legal expertise to assess AI tools against a structured risk framework — evaluating privacy, security, pedagogical value, and algorithmic accountability — before deployment. Most educational institutions currently lack this expertise, creating a significant capacity gap that vendors are filling with their own assessments, which may not be independent or rigorous.
The European Contrast: GDPR and the AI Act in Education
The European regulatory framework provides stronger baseline protections for student data than FERPA, but also faces significant implementation challenges in the AI era. The GDPR's principles of purpose limitation, data minimisation, and storage limitation apply to AI systems deployed in educational settings, and the requirement for a lawful basis for processing — typically consent or legitimate interest — constrains the ways in which student data can be used for AI training or model improvement.
The EU AI Act adds a further layer of obligation for AI systems deployed in educational settings. AI systems used for student assessment, monitoring, or behavioural analysis may be classified as "high-risk" under the Act, triggering requirements for conformity assessment, transparency, human oversight, and registration in the EU database of high-risk AI systems. These requirements are more demanding than anything in the US regulatory framework, but their implementation in educational settings is still in early stages.
The practical challenge for European educational institutions is that many of the AI tools they are deploying are developed by US companies operating under US legal assumptions. Ensuring that these tools comply with GDPR and the AI Act requires contractual provisions, technical audits, and ongoing monitoring that most institutions are not yet equipped to provide.
What Genuine Data Sovereignty in Education Requires
Contractual Foundations
The shift from data residency to genuine data sovereignty in education requires transparency about how AI models are trained, where prompts are processed, and who maintains access to the data lifecycle — questions that most EdTech contracts do not answer.
The Data Processing Agreement (DPA) is the primary legal mechanism through which educational institutions can establish data sovereignty over AI tools. Effective DPAs for AI in education must include explicit prohibitions on model training using student data, requirements for data deletion upon contract termination, breach notification timelines, audit rights, and clear specifications of the purposes for which student data may be processed.
Most existing DPAs in educational settings were not designed with AI-specific risks in mind. Institutions should review and update their DPA templates to address model training, algorithmic profiling, and the use of AI-generated inferences in decision-making.
Centralised Governance and Procurement
Shadow AI — teachers and students using unauthorised consumer-grade tools outside district oversight — is not a marginal phenomenon. It is the dominant mode of AI adoption in education, and it is systematically bypassing every data protection framework institutions have built.
The shadow AI problem cannot be addressed through contractual mechanisms alone — it requires institutional governance structures that make authorised AI tools accessible and attractive enough that teachers and students do not need to resort to unauthorised alternatives. This means centralised procurement processes that evaluate AI tools against privacy, security, and pedagogical criteria; clear communication to staff and students about which tools are authorised and why; and accessible support for using authorised tools effectively.
Institutions should move away from self-service AI tool adoption towards formal, centralised assessment processes. This is not about restricting innovation — it is about ensuring that the tools deployed in educational settings meet the standards of care that students and families have a right to expect.
Technical Controls and Audit Capabilities
Contractual provisions are only as effective as the technical controls and audit capabilities that back them up. Institutions need the technical capacity to monitor how AI tools are using student data, to detect unauthorised data flows, and to audit vendor compliance with contractual obligations. Building this capacity requires investment in technical expertise that most educational institutions currently lack.
Data minimisation — ensuring that AI tools collect only the minimum information necessary for their educational purpose — is both a legal requirement under GDPR and a practical risk management strategy. Institutions should require vendors to demonstrate that their data collection practices are proportionate to their educational function, and should resist the deployment of AI tools that collect broad behavioural data without a clear educational justification.
Conclusion: The Urgency of Educational Data Sovereignty
The deployment of AI in education is accelerating, driven by genuine potential benefits — personalised learning, reduced administrative burden, improved accessibility — and by commercial pressures from a rapidly growing EdTech industry. The governance frameworks needed to ensure that this deployment respects student privacy, prevents algorithmic discrimination, and maintains institutional control over student data are not keeping pace.
The gap between the pace of AI deployment and the maturity of governance frameworks is not merely a technical or legal problem — it is a question of institutional values. Educational institutions exist to serve students, and that service obligation extends to protecting students' data, their privacy, and their right to an education that is not mediated by opaque algorithmic systems that they cannot understand, challenge, or control.
Genuine data sovereignty in education requires more than data residency requirements and FERPA compliance checklists. It requires transparency about how AI systems work, contractual and technical controls over the entire data lifecycle, governance structures that give students and families meaningful agency over their data, and the institutional capacity to audit and enforce these requirements. Building that capacity is one of the most urgent governance challenges in education today — and the window for getting it right, before AI becomes so deeply embedded in educational infrastructure that reform becomes prohibitively difficult, is narrowing.





