- Level
- Undergraduate and graduate (cross-listed)
- Prerequisites
- None required; prior programming experience in Python is strongly recommended
- Length
- 15 weeks, three 50-minute meetings a week
- Tools
- Python, Jupyter, requests, BeautifulSoup, selenium, pypdf, pdfplumber, pandas, Git, GitHub, GitHub Actions
- Textbook
- Web Data Science (open, online)
- Taught
- 3 terms, Fall 2023 to Fall 2026
- License
- CC BY-NC-SA 4.0
Overview
Web Data Science is a course in which the students also edit the textbook. The class reads one chapter of the open textbook Web Data Science each week and proposes improvements to it as GitHub pull requests. Along the way, students learn to retrieve, parse, and analyze web data with Python: the law and ethics of web data access, data formats and web protocols, extraction from web pages and PDFs, and collection from popular APIs. The web makes many kinds of information easy to reach, and a data scientist who can turn web documents and APIs into data can use that information in common research designs.
I teach it as a cross-listed course for upper-division undergraduates (INFO 4617) and graduate students (INFO 5617). There are no required prerequisites, but students should have some experience with Python (loops, functions, lists, and dictionaries). There is no midterm or final exam. A week has three meetings. The first introduces a concept and its companion notebook, and the second is a notebook lab. In the third, a workshop, students propose revisions to the textbook and review their classmates’ proposals.
The four modules (Foundations, Documents, APIs, and Practice) follow the four parts of the textbook, and students finish with an original data collection project. The last exercise in each chapter of the textbook is a graduate extension for INFO 5617 students, which pairs a scholarly reading with a more open-ended task.
Learning objectives
- Explain the history and infrastructure of web data access, and the law and ethics that apply to it.
- Read and parse common web data formats such as XML and JSON.
- Retrieve data from HTML and PDF documents, and automate the extraction.
- Access popular APIs to collect data for common research designs.
- Use version control and open-source workflows (Git and GitHub pull requests) to collaborate on a shared codebase.
- Document, communicate, and critically evaluate data collection code and technical writing.
Topics
Each week covers one chapter of the textbook. Students read the chapter, work through its companion notebook in the lab, and propose a change to the chapter in the workshop.
| Week | Module | Topic |
|---|---|---|
| 1 | Foundations | Introduction and setup: the Python environment, Git, GitHub, and the textbook revision workflow |
| 2 | Foundations | Ethics and law: terms of service, copyright, privacy, and research ethics |
| 3 | Foundations | The post-API age: changing access to data and platform transparency |
| 4 | Foundations | Data formats: XML and JSON and the libraries that parse them |
| 5 | Foundations | Protocols: web architecture, IP, and HTTP with urllib and requests |
| 6 | Documents | Static web pages: HTML elements and parsing with BeautifulSoup |
| 7 | Documents | Archived web pages: the Internet Archive’s Wayback Machine |
| 8 | Documents | Dynamic web pages: rendering and scraping pages with selenium |
| 9 | Documents | PDFs: extracting text and tables with pypdf and pdfplumber |
| 10 | APIs | Wikipedia: an introduction to APIs with Wikipedia’s APIs |
| 11 | APIs | Government data: the Census, Federal Reserve (FRED), and FEC APIs |
| 12 | APIs | Social and media platforms: content, social graphs, and activity streams |
| 13 | APIs | AI and language models: using large language models through their APIs |
| 14 | Practice | Automation: scheduled data collection with GitHub Actions |
| 15 | Practice | Research design and final projects: matching web data sources with research designs; presentations |
Readings
The main reading is my open textbook Web Data Science, one chapter for each week above. Each chapter has learning objectives, a guided tutorial, exercises, a section on social history and the public interest, common debugging problems, and further reading.
- Foundations: Introduction to Web Data Science; Ethics, Law, and Responsible Data Collection; The Post-API Age; Data Formats: XML and JSON; Web Architecture and Protocols
- Documents: Parsing Static Web Pages; Archived Web Pages and the Wayback Machine; Dynamic Web Pages with Selenium; Extracting Data from PDFs
- APIs: Introduction to APIs: Wikipedia; Government Data APIs; Social and Media Platform APIs; AI and Language Model APIs
- Practice: Automating Data Collection; Research Design with Web Data
From the further reading in the chapters, these works are open or have a DOI:
- Foundations: Fiesler, Beard, and Keegan, “No Robots, Spiders, or Scrapers” (2020); Freelon, “Computational Research in the Post-API Age” (2018); Bruns, “After the ‘APIcalypse’” (2019); Perriam, Birkbak, and Freeman, “Digital Methods in a Post-API Environment” (2020); Davidson et al., “Platform-Controlled Social Media APIs Threaten Open Science” (2023)
- Documents: Arora et al., “Using the Wayback Machine to Mine Websites in the Social Sciences” (2016)
- APIs: Mesgari et al., “The Sum of All Human Knowledge” (2015); Baumgartner et al., “The Pushshift Reddit Dataset” (2020); Diener and Delcourt, “Mastodon.py” (2026); Ziems et al., “Can Large Language Models Transform Computational Social Science?” (2024)
- Practice: Wilson et al., “Good Enough Practices in Scientific Computing” (2017); Gebru et al., “Datasheets for Datasets” (2021); Lazer et al., “Computational Social Science: Obstacles and Opportunities” (2020)
Assignments and assessment
| Assignment | Share of grade |
|---|---|
| Notebook labs | 30% |
| Textbook revisions | 15% |
| Attendance | 15% |
| Final project | 40% |
Notebook labs happen in class every week. Students work through the chapter’s companion notebook together and share their implementation. Labs are graded on participation and completion, and the code does not have to be perfect. The two lowest lab scores are dropped.
Textbook revisions are also weekly. Each student proposes a change to the textbook (a correction, a clarification, a new example, an exercise, or a better explanation) as an issue or a pull request to the book’s repository, and reviews two classmates’ pull requests. A 10-point rubric scores whether the problem is located, evidenced, and actionable, and how good the reviews are. The grade is for the proposal and the review, so a change does not have to be merged to earn credit. The revision framework and the pull request walkthrough explain the process.
Attendance is required. The methods build on each other from week to week, and the labs and workshops happen in class.
The final project is a portfolio piece, done alone or in pairs, that matches a web data source with a research design. It has three graded parts: a proposal in week 10, a presentation in the final week, and a project repository with a write-up.
Adopt this course
Students use the Anaconda distribution of Python with Jupyter, plus Git and a free GitHub account. Chapter 1 creates one conda environment with every library the book uses, and the course’s setup handout walks students through Anaconda, Git, and GitHub before the first lab. Chapter 8 also installs a browser for Playwright, and Chapters 11 to 13 need students’ own API keys. For general computing skills (Jupyter, debugging, regular expressions, version control, and secrets management), the book points students to the Missing Manual for Information Scientists.
Most chapters collect live data from the web and from the APIs of Wikipedia, the U.S. Census Bureau, FRED, the FEC, Reddit, Spotify, Bluesky, Mastodon, OpenAI, and Anthropic. For the chapters that read fixed files, the book’s data folder keeps stable copies: an English stopword list (Chapter 7), and City of Boulder revenue reports and city council minutes as PDFs, with a manifest of their sources (Chapter 9).
The book’s notebooks folder has one companion notebook per chapter. The notebooks are generated from the chapters and left unexecuted on purpose, so that students run the code and fix the errors themselves. The Fall 2026 repository has lecture slides for weeks 1 to 14 (LaTeX and PDF) and handouts for the revision workshops. The Fall 2024 repository has a lecture notebook for each week of the earlier design.
The topics fill a 15-week semester with three meetings a week, plus one week off for a break. Earlier terms taught a similar sequence in two 75-minute meetings a week. To teach the course without the textbook revisions, replace the workshops with module assignments, as the Fall 2023 and Fall 2024 terms did. The textbook has its own CC BY-NC-SA 4.0 license, and the Fall 2024 and Fall 2026 course repositories are under the MIT License; the material on this page is under CC BY-NC-SA 4.0.
Past offerings
The course ran twice as a special topics course (INFO 4871/5871) before it got its own number, INFO 4617, in Fall 2026. For that term I rebuilt it around the open textbook.
| Term | Design | Materials |
|---|---|---|
| Fall 2023 (INFO 4871/5871) | Three modules (fundamentals, documents, APIs), with APIs for Wikipedia, the Census, Mastodon, and Reddit, and automation with scrapy and GitHub Actions; graded on attendance (15%), module assignments (60%), and a final project (25%) | — |
| Fall 2024 (INFO 4871) | The 2023 design, plus weeks on the post-API age and on AI APIs | Repository |
| Fall 2026 (INFO 4617) | The design on this page: one textbook chapter a week, with notebook labs and textbook revisions | Repository |
Course materials on this page are licensed under CC BY-NC-SA 4.0: reuse and adapt them with credit, not for commercial use, and share adaptations under the same license. Last reviewed October 8, 2026.