- Level
- Undergraduate and graduate (cross-listed as INFO 4871 and INFO 5871)
- Prerequisites
- Prior coursework in data science, statistics, or an information or media studies methods course; comfort with Python and pandas
- Length
- 15 weeks, two 75-minute meetings a week
- Tools
- Python, Jupyter, pandas, requests, Beautiful Soup, pdfplumber, scikit-learn, fairlearn, folktables, pytest, SQLite, git, GitHub, Quarto, Zenodo
- Textbook
- Public Interest Data Science (open, online)
- Taught
- First offered Spring 2027
- License
- CC BY-NC-SA 4.0
Overview
Public Interest Data Science asks how democratic societies can observe, audit, and hold accountable the data-intensive systems that now affect housing, labor, policing, public health, elections, and the environment, and what data scientists can do about it. I designed the course, and it has not been taught yet. The first offering is in Spring 2027, so this page describes the planned design.
The course is built on a framework I call public interest data infrastructuring. Three pressures narrow the public’s capacity to see and govern data systems: enclosure (who can observe?), exemption (who must answer?), and erosion (what can be remembered?). Three values push back: openness, oversight, and ownership. Those values last only when someone builds them into an installed base of identifiers, records, access tiers, retention rules, and pathways to remedy. The course also treats “the public interest” itself as a contested and partial idea. Each week will pair a framing reading with a critical reading, and a US case with a non-US counter-case.
There are four three-week modules, plus an opening week and a closing week. Each module borrows the methods of one profession (its lineage) and works at one level of government: journalism at the city, planning (with civic tech and refusal) at the county, law at the state, and assurance (accounting and engineering) at the nation. Each module will end in a portfolio piece, which is a technical artifact plus a public text in the genre that the profession uses to reach the public (an op-ed, a public comment, testimony, and a report). Students will also contribute a chapter-scale revision to the course’s open textbook. There is no midterm or final exam. INFO 4871 (undergraduate) and INFO 5871 (graduate) are special-topics course numbers. The two levels will meet together, and graduate students will write an extra methods memo for each piece.
Learning objectives
- Diagnose a data-governance problem in terms of enclosure, exemption, and erosion, and identify which elements of the installed base (linkability, interpretability, continuity, safe scrutiny, authority, and remedy) are missing.
- Borrow the methods of a public interest profession (journalism, law, accounting, engineering, or planning) and name the limits of that borrowing.
- Execute the technical work of public interest data science: reproduce a published analysis, scrape responsibly, request and extract public records, run and test a disaggregated audit, audit link rot, document a dataset, and design tiered access or a refusal specification.
- Write for non-academic audiences in four genres: op-ed, public comment, testimony, and report.
- Work across levels of government, and match the evidence, the genre, and the request to the city, county, state, or national body that can act on it.
- Critique your own interventions: whose interests they serve, whom they could harm, which communities were not consulted, and where the framework itself fails.
- Situate each piece in the research literature and defend its methodological choices against that literature (graduate students only).
Topics
The first meeting each week will be for ideas: a lecture and a case discussion of the week’s paired readings. The second meeting is a lab that follows a tutorial in the textbook, or a genre workshop. The last week of each module ends with a studio for the portfolio piece. Every module draws on three cases: a national anchor case, a non-US counter-case, and a local case at the module’s level of government.
| Week | Module | Topic |
|---|---|---|
| 1 | Opening | The public interest: who gets to define it; its emancipatory and exclusionary histories |
| 2 | Opening | Infrastructuring: the three pressures, the three values, and the installed base. Lab: a first API request and a first pull request |
| 3 | Journalism (city) | Enclosure and the watchdog: API shutdowns, paywalls, verification, and provenance. Lab: reproduce ProPublica’s “Machine Bias” |
| 4 | Journalism (city) | Openness: posting is not publishing; licensing as governance. Lab: a responsible scraper and a city dataset |
| 5 | Journalism (city) | The op-ed: hook, argument, ask, and pitch. Studio: Piece 1 |
| 6 | Planning (county) | Erosion: link rot, dataset deprecation, retention schedules, and advocacy planning. Lab: an erosion audit of county records |
| 7 | Planning (county) | Ownership: stewardship, Indigenous data sovereignty, data justice, and refusal. Lab: a stewardship plan or a refusal specification |
| 8 | Planning (county) | The public comment: reading a staff report or draft plan; who is heard. Studio: Piece 2 |
| 9 | Law (state) | Exemption: legal carve-outs, trade secrecy, terms of service, and the Computer Fraud and Abuse Act. Lab: an evidence ledger and the state open records statute |
| 10 | Law (state) | Oversight as remedy: routine auditability, escalation, and enforcement. Lab: from state documents to a table |
| 11 | Law (state) | Testimony: written and oral testimony to a state legislative committee. Studio: Piece 3 |
| 12 | Assurance (nation) | Accounting and audit-washing: independence, standards, and disclosure. Lab: a disaggregated audit |
| 13 | Assurance (nation) | Engineering and procurement: safety codes, incident review, worker observatories, and civic tech. Lab: test the audit |
| 14 | Assurance (nation) | The report: executive summary, methods, findings, and recommendations. Studio: Piece 4 |
| 15 | Closing | The limits of the framework: critiques, US-centrism, and records over refusal; archiving with a DOI; book showcase |
Readings
I wrote the open textbook Public Interest Data Science for this course, and students will read and revise it together. Each chapter combines concepts with worked examples and exercises. The course assigns the chapters by module, so they come out of book order.
- Opening: Introduction to Public Interest Data Science; The Installed Base; Public Interest Technology and Data for Good; Washington and Cheung, “Towards Defining the Public Interest in Technology” (2024); Stapleton et al., “Who Has an Interest in ‘Public Interest Technology’?” (2022); Plantin et al., “Infrastructure Studies Meet Platform Studies” (2016); Aula and Bowles, “Stepping Back from Data and AI for Good” (2023)
- Journalism (city): Enclosure and Openness; Journalism; Op-Eds; Bruns, “After the ‘APIcalypse’” (2019); Zong and Matias, “Data Refusal from Below” (2024); Angwin et al., “Machine Bias” (2016); Flores, Bechtel, and Lowenkamp, “False Positives, False Negatives, and False Analyses” (2016)
- Planning (county): Erosion and Ownership; Planning; Critical Data Studies, Data Justice, and Refusal; Public Comments; Davidoff, “Advocacy and Pluralism in Planning” (1965); Denton et al., “On the Genealogy of Machine Learning Datasets” (2021); Gebru et al., “Datasheets for Datasets” (2021); Gaffney and Matias, “Caveat Emptor, Computational Social Science” (2018); D’Ignazio and Klein, Data Feminism (2020); Carroll et al., “The CARE Principles for Indigenous Data Governance” (2020); Taylor, “What Is Data Justice?” (2017)
- Law (state): Exemption and Oversight; Law; Testimony; Albiston, “Public Interest Law Organizations and the Two-Tier System of Access to Justice” (2017); Warren et al., “RequestAtlas: Supporting the Slow and Iterative Process of Requesting Public Records” (2025); Amnesty International, “Xenophobic Machines” (2021)
- Assurance (nation): Accounting and Auditing; Engineering; GovTech and Civic Tech; Reports; Raji et al., “Closing the AI Accountability Gap” (2020); Goodman and Trehu, “AI Audit-Washing and Accountability” (2023); Buolamwini and Gebru, “Gender Shades” (2018); Baker, “What Is the Meaning of ‘the Public Interest’?” (2005); Miceli, Posada, and Yang, “Studying Up Machine Learning Data” (2022); Bharosa, “The Rise of GovTech” (2022); Taylor, “Public Actors Without Public Values” (2021); Gray and Bounegru (eds.), The Data Journalism Handbook (2019)
- Closing: Archival Deposits; The Limits of the Framework; graduate students also read Research Proposals
The course also uses my paper that defines the framework, Keegan (2026), “Public interest data infrastructuring.” I have not yet chosen some of the critical readings for the genre weeks.
Assignments and assessment
| Assignment | Share of grade |
|---|---|
| Participation, including a textbook contribution | 15% |
| Portfolio pieces (4 × 15%) | 60% |
| Final project | 25% |
Participation covers attendance, discussion, and the weekly labs (graded complete or incomplete). It also includes a required contribution to the textbook. Each student will propose one chapter-scale revision in an issue, open a pull request, get it merged, present it at an end-of-term showcase, and give substantive reviews on at least three classmates’ pull requests. A chapter-scale revision is a new or rewritten section with narrative, a runnable tutorial, verified references, and curated resources, following the book’s contribution guide. Typo fixes do not meet the requirement.
Portfolio pieces come at the end of each module. Every piece has four parts: (1) a technical artifact, which is a reproducible notebook or small repository that a classmate can clone and run; (2) a public text in the module’s genre for a named audience; (3) an installed-base note of about 300 words on the pressure, the installed-base elements built or still missing, provenance, license, and AI use; and (4) for graduate students, a methods memo of 750 to 1,000 words that cites at least five scholarly sources. Each piece is scored out of 100: technical soundness and reproducibility (30), the public text (30), the installed-base note (20), and critical framing (20). The four pieces are:
- Piece 1 (journalism, city): reproduce one published finding and extend it with responsibly collected city data and a provenance file; write a 750-word op-ed for a named local outlet, with a pitch.
- Piece 2 (planning, county): audit the erosion of county records or datasets (link rot, Wayback Machine recovery, and a datasheet), using records requests that the county has already answered; write a stewardship plan or a refusal specification; write a public comment of about 1,000 to 1,500 words on a live county planning or land use matter.
- Piece 3 (law, state): build an evidence ledger for a deployed system, turn state documents or released records into a table, and map each exemption and remedy to the state law behind it; write about 1,000 words of testimony and record a 3- to 5-minute oral statement for a state legislative committee.
- Piece 4 (assurance, nation): run a disaggregated audit on national data, with a manifest, thresholds, a changelog, and tests in continuous integration; write a report section of about 2,000 words for a named federal body.
The final project takes one piece further. Students will deepen the analysis (new data, a second level of government, or a comparison with the non-US counter-case), recast it in a different genre (graduate students may write a research proposal), archive the artifact and text in a repository that mints a DOI (such as Zenodo), and write a 500-word critical-reflection memo. With the instructor’s approval, a student can file a new public records request for the final project.
Adopt this course
Students will use Python 3.11 or newer (through Anaconda or uv) with Jupyter, git and GitHub for version control and peer review, Quarto for documents and the textbook, and SQLite. The libraries are requests, beautifulsoup4, pandas, pdfplumber, scikit-learn, fairlearn, folktables, pytest, and the internetarchive client. Each student keeps a portfolio in their own GitHub repository, and the Zenodo sandbox is a place to practice deposits before minting a real DOI.
The data come from public sources at each level of government. Examples are the Wikimedia pageviews API; ProPublica’s COMPAS data from “Machine Bias”; city open data portals and council records (City of Boulder and City and County of Denver); Boulder County’s archive of answered records requests, planning documents, and meeting records; Colorado bill text, fiscal notes, and agency reports; and American Community Survey data through folktables. To teach the course somewhere else, replace the local case in each module with a city, county, state, and national body where you teach. The national anchor cases and the non-US counter-cases do not depend on Colorado.
The public materials are the textbook and its source repository. The labs follow tutorials in the chapters, and each genre chapter ends with an artifact that students can submit. The course slides, lab handouts, and assignment handouts are not public.
The schedule fills 15 weeks of class plus a break week. Each module takes three weeks: two weeks of ideas and labs, then a genre week with a studio. In the first offering, the law module will span spring break, and its testimony week will fall while the Colorado General Assembly is in session. If you adopt the course, time the public comment and testimony weeks to live county matters and legislative hearings. The textbook’s text is under CC BY 4.0 and its code under the MIT License; the material on this page is under CC BY-NC-SA 4.0.
Past offerings
The course has not been taught yet. Spring 2027 will be its first term.
| Term | Design | Materials |
|---|---|---|
| Spring 2027 | First offering (the design on this page) | — |
Course materials on this page are licensed under CC BY-NC-SA 4.0: reuse and adapt them with credit, not for commercial use, and share adaptations under the same license. Last reviewed October 8, 2026.