INFO 4617/5617

Web Data Science

Level
Undergraduate and graduate (cross-listed)
Prerequisites
None required; prior programming experience in Python is strongly recommended
Length
15 weeks, three 50-minute meetings a week
Tools
Python, Jupyter, requests, BeautifulSoup, selenium, pypdf, pdfplumber, pandas, Git, GitHub, GitHub Actions
Textbook
Web Data Science (open, online)
Taught
3 terms, Fall 2023 to Fall 2026
License
CC BY-NC-SA 4.0

Overview

Web Data Science is a course in which the students also edit the textbook. The class reads one chapter of the open textbook Web Data Science each week and proposes improvements to it as GitHub pull requests. Along the way, students learn to retrieve, parse, and analyze web data with Python: the law and ethics of web data access, data formats and web protocols, extraction from web pages and PDFs, and collection from popular APIs. The web makes many kinds of information easy to reach, and a data scientist who can turn web documents and APIs into data can use that information in common research designs.

I teach it as a cross-listed course for upper-division undergraduates (INFO 4617) and graduate students (INFO 5617). There are no required prerequisites, but students should have some experience with Python (loops, functions, lists, and dictionaries). There is no midterm or final exam. A week has three meetings. The first introduces a concept and its companion notebook, and the second is a notebook lab. In the third, a workshop, students propose revisions to the textbook and review their classmates’ proposals.

The four modules (Foundations, Documents, APIs, and Practice) follow the four parts of the textbook, and students finish with an original data collection project. The last exercise in each chapter of the textbook is a graduate extension for INFO 5617 students, which pairs a scholarly reading with a more open-ended task.

Learning objectives

  1. Explain the history and infrastructure of web data access, and the law and ethics that apply to it.
  2. Read and parse common web data formats such as XML and JSON.
  3. Retrieve data from HTML and PDF documents, and automate the extraction.
  4. Access popular APIs to collect data for common research designs.
  5. Use version control and open-source workflows (Git and GitHub pull requests) to collaborate on a shared codebase.
  6. Document, communicate, and critically evaluate data collection code and technical writing.

Topics

Each week covers one chapter of the textbook. Students read the chapter, work through its companion notebook in the lab, and propose a change to the chapter in the workshop.

Week Module Topic
1 Foundations Introduction and setup: the Python environment, Git, GitHub, and the textbook revision workflow
2 Foundations Ethics and law: terms of service, copyright, privacy, and research ethics
3 Foundations The post-API age: changing access to data and platform transparency
4 Foundations Data formats: XML and JSON and the libraries that parse them
5 Foundations Protocols: web architecture, IP, and HTTP with urllib and requests
6 Documents Static web pages: HTML elements and parsing with BeautifulSoup
7 Documents Archived web pages: the Internet Archive’s Wayback Machine
8 Documents Dynamic web pages: rendering and scraping pages with selenium
9 Documents PDFs: extracting text and tables with pypdf and pdfplumber
10 APIs Wikipedia: an introduction to APIs with Wikipedia’s APIs
11 APIs Government data: the Census, Federal Reserve (FRED), and FEC APIs
12 APIs Social and media platforms: content, social graphs, and activity streams
13 APIs AI and language models: using large language models through their APIs
14 Practice Automation: scheduled data collection with GitHub Actions
15 Practice Research design and final projects: matching web data sources with research designs; presentations

Readings

The main reading is my open textbook Web Data Science, one chapter for each week above. Each chapter has learning objectives, a guided tutorial, exercises, a section on social history and the public interest, common debugging problems, and further reading.

From the further reading in the chapters, these works are open or have a DOI:

Assignments and assessment

Assignment Share of grade
Notebook labs 30%
Textbook revisions 15%
Attendance 15%
Final project 40%

Notebook labs happen in class every week. Students work through the chapter’s companion notebook together and share their implementation. Labs are graded on participation and completion, and the code does not have to be perfect. The two lowest lab scores are dropped.

Textbook revisions are also weekly. Each student proposes a change to the textbook (a correction, a clarification, a new example, an exercise, or a better explanation) as an issue or a pull request to the book’s repository, and reviews two classmates’ pull requests. A 10-point rubric scores whether the problem is located, evidenced, and actionable, and how good the reviews are. The grade is for the proposal and the review, so a change does not have to be merged to earn credit. The revision framework and the pull request walkthrough explain the process.

Attendance is required. The methods build on each other from week to week, and the labs and workshops happen in class.

The final project is a portfolio piece, done alone or in pairs, that matches a web data source with a research design. It has three graded parts: a proposal in week 10, a presentation in the final week, and a project repository with a write-up.

Adopt this course

Students use the Anaconda distribution of Python with Jupyter, plus Git and a free GitHub account. Chapter 1 creates one conda environment with every library the book uses, and the course’s setup handout walks students through Anaconda, Git, and GitHub before the first lab. Chapter 8 also installs a browser for Playwright, and Chapters 11 to 13 need students’ own API keys. For general computing skills (Jupyter, debugging, regular expressions, version control, and secrets management), the book points students to the Missing Manual for Information Scientists.

Most chapters collect live data from the web and from the APIs of Wikipedia, the U.S. Census Bureau, FRED, the FEC, Reddit, Spotify, Bluesky, Mastodon, OpenAI, and Anthropic. For the chapters that read fixed files, the book’s data folder keeps stable copies: an English stopword list (Chapter 7), and City of Boulder revenue reports and city council minutes as PDFs, with a manifest of their sources (Chapter 9).

The book’s notebooks folder has one companion notebook per chapter. The notebooks are generated from the chapters and left unexecuted on purpose, so that students run the code and fix the errors themselves. The Fall 2026 repository has lecture slides for weeks 1 to 14 (LaTeX and PDF) and handouts for the revision workshops. The Fall 2024 repository has a lecture notebook for each week of the earlier design.

The topics fill a 15-week semester with three meetings a week, plus one week off for a break. Earlier terms taught a similar sequence in two 75-minute meetings a week. To teach the course without the textbook revisions, replace the workshops with module assignments, as the Fall 2023 and Fall 2024 terms did. The textbook has its own CC BY-NC-SA 4.0 license, and the Fall 2024 and Fall 2026 course repositories are under the MIT License; the material on this page is under CC BY-NC-SA 4.0.

Past offerings

The course ran twice as a special topics course (INFO 4871/5871) before it got its own number, INFO 4617, in Fall 2026. For that term I rebuilt it around the open textbook.

Term Design Materials
Fall 2023 (INFO 4871/5871) Three modules (fundamentals, documents, APIs), with APIs for Wikipedia, the Census, Mastodon, and Reddit, and automation with scrapy and GitHub Actions; graded on attendance (15%), module assignments (60%), and a final project (25%) —
Fall 2024 (INFO 4871) The 2023 design, plus weeks on the post-API age and on AI APIs Repository
Fall 2026 (INFO 4617) The design on this page: one textbook chapter a week, with notebook labs and textbook revisions Repository

Course materials on this page are licensed under CC BY-NC-SA 4.0: reuse and adapt them with credit, not for commercial use, and share adaptations under the same license. Last reviewed October 8, 2026.