Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
Design and Policy-1: Shadow Archives in the Age of Extraction: Infrastructural Citizenship and the Technopolitics of AI Training Datasets
| ||
| Presentations | ||
Shadow Archives in the Age of Extraction: Infrastructural Citizenship and the Technopolitics of AI Training Datasets Kennesaw State University, United States of America Research Question: As Large Language Models (LLMs) increasingly extract data from extra-legal digital repositories for algorithmic training, how do shadow libraries actively negotiate their institutional legitimacy, and how do they assert infrastructural agency against commercial AI developers engaging in mass data enclosure? Rather than treating these platforms as passive repositories of illicit content, this research interrogates the political economy of AI training data by centering shadow archives as active sociotechnical institutions navigating a rapidly shifting landscape of corporate extraction, copyright litigation, and emergent governance frameworks. Methodology: This research applies digital infrastructure ethnography and platform hermeneutics to trace the institutional evolution of Anna's Archive, a prominent shadow library aggregating over 1.3 billion bibliographic records. The primary dataset consists of public-facing communications, technical documentation, architectural change logs, and metadata schemas collected between early 2023 and early 2026, capturing a critical period of platform transformation driven by the rise of large-scale AI data extraction. The analysis maps the platform's key architectural shifts, focusing on its implementation of SFTP access management protocols designed to throttle and selectively permit scraping loads from commercial AI developers. Secondary analysis draws on publicly available litigation documents from copyright infringement cases involving datasets derived from shadow library content, including the Books3 dataset used in training Meta's LLaMA models, to triangulate the legal pressures shaping platform behavior. Analytically, the research employs Mark Suchman's (1995) framework of organizational legitimacy to systematically code the archive's tactical and discursive responses to corporate extraction. This coding traces the platform's evolution from a human-centric reading resource operating under a piracy frame to a self-described infrastructural actor managing an algorithmic data supply chain, revealing how legitimacy is actively constructed under conditions of legal precarity. Disciplinary Fields: Critical Data Studies, Science and Technology Studies (STS), and Information Policy. Novelty & Policy Relevance: AI policy debates remain heavily concentrated on copyright infringement litigation and the regulation of generative outputs, scrutinizing what AI systems produce rather than what they consume. This research redirects analytical attention toward the shadow supply chains that materially power foundational models, examining the infrastructural conditions of AI training data at the point of extraction. Tracing the historical genealogy of shadow libraries from Soviet-era samizdat through early digital piracy networks to their present incarnation as de facto data utilities contextualizes these platforms as evolving architectures of epistemic disobedience with deep institutional histories. This historical grounding addresses a persistent gap in AI governance scholarship, which tends to treat training data extraction as a novel legal problem rather than the latest iteration of a long-standing tension between knowledge enclosure and open access. These findings bear directly on several pending regulatory instruments, including the EU AI Act's Article 53 training data disclosure requirements and proposed U.S. copyright transparency mandates currently before Congress, all of which lack analytical frameworks for evaluating non-commercial infrastructural actors operating outside formal intellectual property regimes. Expected Results: Shadow archives have transitioned from clandestine repositories into active sociotechnical agents that mobilize their scale to force negotiations with commercial AI developers. Platforms like Anna's Archive have strategically embraced their role as data intermediaries, implementing tiered access systems and forging selective partnerships with AI developers including DeepSeek, blurring the boundary between illicit archive and legitimate infrastructure provider. Shadow libraries construct pragmatic legitimacy through demonstrated technical competence and indispensable scale, while simultaneously constructing moral legitimacy by deploying anti-enclosure narratives that frame open knowledge access as a fundamental human right imperiled by corporate data hoarding. Regulatory models that treat AI training data exclusively as a vector for intellectual property theft fail to account for the infrastructural citizenship these archives exhibit. By positioning themselves as custodians of the global knowledge commons, shadow libraries are actively shaping the normative terrain of AI governance from outside formal policy channels. Future AI governance frameworks should incorporate data provenance transparency mandates that acknowledge the non-commercial maintenance labor inherent in shadow infrastructures, and policymakers need to engage seriously with the material realities of how foundational models are actually built.
| ||
