Applied AI for Public Policy

← 12-week course plan

You are on a small screen. This page is designed for a laptop, with three columns and hands-on steps appearing beside the content. On a phone it collapses to one column and the hands-on steps move to the end of the page. Reading here is fine; for the exercises you need your laptop anyway.

Week 4 · Oct 1, 2026 · Thursday, 9:00–11:30 am · 280 Brook Street, Room 110

AI for data collection

Collect a documented dataset and explain how its coverage supports or changes the project question.

Opening

Today

  • Why a policy analyst would collect data.
  • APIs, with the Federal Register as the example.
  • Sources without an API: web scraping.
  • Collecting from several sources at once.
  • What your data leave out.

Why collect data at all

  • Policy analysts usually work from tables that someone else published.
  • Last week's reading said the same. Most policy research reuses what others produced.
  • A published table may not fit your question. It may have the wrong years, the wrong places, or too little detail.
  • With an agent, you can collect the documents your question needs in one class session.
  • So the question changes. Are these data worth collecting, and what do they leave out?

Readings

Datasheets for Datasets

APIs · do it on the right

What an API is

  • A government website is made for a person who reads one page at a time. A policy analysis often needs hundreds of documents at once.
  • An agency can also offer a second way to get the same documents, made for programs. This is an API, short for application programming interface.
  • A program sends the API a request. Parameters in the request say which documents you want, such as the agency, the dates, and the document type.
  • The documents come back with no page layout, in a form a script can read directly. Every document has the same fields, such as title and date.
  • Some APIs require a key. A key identifies you and limits how fast you may send requests.
  • Keep your key out of Git. The Federal Register API needs no key.
  • Many federal agencies offer an API. Many state agencies and city governments do not. Sources without an API covers what to do then.

The Federal Register

  • The Federal Register is the U.S. government's daily journal.
  • Every federal agency publishes its proposed rules, final rules, and notices there.
  • A proposed rule is a draft that the public can comment on. A final rule is the version that takes effect.
  • A proposed rule usually lists a docket number and the date public comments close.
  • A policy question it can answer: what did the Environmental Protection Agency propose in 2025, and when could the public respond?
  • In 2025 the EPA published 242 proposed rules. On the website, that is 242 pages to open one at a time.

Start on the website

  • Open federalregister.gov. Search for proposed rules from the EPA in 2025.
  • Open one document. Find the title, the agency, the date, the summary, the docket number, and the comment deadline.
  • The API returns those same pieces for every matching document, with no page layout.
  • The website is for people. When a script asks the website for a page, it is sent to a block page.
  • The API answers the same script.

The API documentation

  • The Federal Register API documentation lists what you can ask for and the name of every parameter.
  • Agency names must be spelled the way the API spells them.
  • In the URL, each agency is written in lowercase with hyphens. The EPA is environmental-protection-agency. The Department of Education is education-department.
  • The agency list gives this form for every agency, in the field labeled slug.
  • Read the documentation before you ask the agent for anything. You need it to check the URL the agent builds.

Two example requests

  • Proposed rules from the Environmental Protection Agency, published in 2025:
https://www.federalregister.gov/api/v1/documents.json?conditions[agencies][]=environmental-protection-agency&conditions[publication_date][year]=2025&conditions[type][]=PRORULE&per_page=5copy
  • Final rules from the Department of Education, published in 2024:
https://www.federalregister.gov/api/v1/documents.json?conditions[agencies][]=education-department&conditions[publication_date][year]=2024&conditions[type][]=RULE&per_page=5copy
  • Everything before the question mark is the endpoint, the address of the API. It is the same in both.
  • Everything after it is the parameters. The two URLs differ in three of them: the agency, the year, and the type.
ParameterWhat it setsValues you can use
conditions[agencies][]the agencythe agency, as written in the agency list: environmental-protection-agency, education-department
conditions[publication_date][year]the year of publicationa four-digit year, such as 2024
conditions[type][]the kind of documentPRORULE for a proposed rule, RULE for a final rule, NOTICE for a notice, PRESDOCU for a presidential document
per_pagehow many documents one page showsa number, such as 5
  • Copy the agency from the list. A guess such as department-of-education returns an error.
  • What comes back is in JSON format. It is built like a Python dictionary: each field name paired with its value.
  • Raw JSON is hard to read in a browser. To see its structure, paste it into a formatter such as jsonformatter.org.

Your own example

  • Choose an agency that works on your policy topic, such as the Labor Department.
  • Open the agency list. Search the page for "Labor". In that agency's entry, the field labeled slug reads labor-department.
  • Start from either example URL above. Put labor-department in the agency parameter.
  • Pick a year. Pick a type from the table above: PRORULE, RULE, or NOTICE.
  • Reload. Run the same search on the website. The two counts should match.
  • You can also describe what you want and ask the agent to build the URL. Read every parameter before you run it.

What you do when the agent writes the request

  • In a programming course, you would read the API documentation and write the code yourself.
  • Here you say what you want in plain words. The agent reads the documentation, writes the code, and runs it.
  • Your job is checking:
    • read the URL the agent built,
    • compare a few documents with the same pages on federalregister.gov,
    • check that the script stays under the source's request limit.
  • This check matters. On a replication benchmark, agents scored about 96 on running an analysis.
  • The same agents scored about 31 on finding the data the analysis needed.

Nguyen et al., "ReplicatorBench" (2026), arXiv:2602.11354.

Every API works the same way

APIYou sendYou get back
Federal Registeran agency, a year, a document typethe matching documents
Regulations.gova docket numberthe public comments on that proposed rule
a language model, such as Claude or GPTa promptthe model's answer
  • You send a request with parameters. Data come back.
  • In Week 5, you send the documents you collect today to the third row.

Keep the raw data

  • Save the response exactly as it came back, in data/. Never edit that file.
  • Next to it, write down:
    • the full request URL,
    • the parameters,
    • the date and time you collected it.
  • Do all cleaning in a script. Then the clean table can always be rebuilt from the raw file.
  • This record of where the data came from is called provenance.
Live demo

Collecting Federal Register documents

  • Collect up to 100 rules and proposed rules that one agency published in 2025.
  • Note how the API chose those 100: the sort order and the filters.
  • Make a table of the documents by month. Label it as the first 100, since the agency may have published more.
  • Change the agency and run it again.
  • Open a few documents on federalregister.gov. Compare the saved title, date, and abstract with the page.

Public sources with an API

SourceWhat is in it
Federal Registerrules, proposed rules, and notices from federal agencies, with abstracts and links
Bureau of Labor Statisticsemployment, wages, and prices, by area and industry
Socratathe platform many city governments use to publish permits, inspections, and service requests
data.govthe federal catalog of datasets, for finding an agency's data
USAspendingfederal contracts, grants, and other awards
EPA and HUDenvironmental monitoring and housing program data

Does your source have an API?

  • Before collecting from any website, check how it publishes its data. Check in this order:
    • a ready-made download, such as a CSV file,
    • an API,
    • the web page itself.
  • A download is the easiest route. The page itself is the last resort.
  • To find an API, search the web for the site's name and "API". Many sites also have a page called "Developers" or "Data".
  • If you find an API, open its documentation. Check what you can ask for and whether you need a key.

Most sources have no API

  • The Federal Register is the easy case. It has an API and needs no key.
  • Many state agencies and city governments offer neither an API nor a download. Their data are only on their web pages.
  • Your project's source may be one of these.
  • The agent can still collect from those pages. The next section shows how.

Sources without an API · do it on the right

Reading a page directly

  • The simplest way is to ask the agent to read the page directly.
    • In Claude Code, the tool is called WebFetch.
    • In Codex, it is web search. Start Codex with codex --search to turn it on.
  • You give it a URL and say what you want, such as "get the monthly unemployment table from this page and save it as a CSV."
  • What the agent gets back can be incomplete. In Claude Code, a second, smaller model reads the page and passes on only its answer. Long pages are cut off first.
  • Every page costs model tokens. In a test for this course, Codex used about 10,000 tokens to read one page.
  • Nothing is saved to rerun. There is no script and no copy of the page.
  • Use it to look at one page or a few.
  • For more pages, or for a collection you will repeat, the agent needs to write a script. That is web scraping.

Claude Code tools reference · Codex CLI documentation

What web scraping is

  • Every web page is built from HTML, the code your browser turns into the page you see.
  • Web scraping is a script that downloads a page's HTML and pulls out the pieces you want.
  • The pieces can be a table, a list of reports, dates, prices, or the text of each page.
  • People scrape when:
    • the site has no API,
    • its API leaves out what they need,
    • the site shows content to readers and offers no way to reuse it.
  • The agent writes the script. Your job is to judge whether a page can be scraped, and to check what comes back.

Is a page easy to scrape?

  • Run four checks in your browser:
    • Can you see the content without logging in?
    • Does the content stay the same when you refresh?
    • Does page 2 look like page 1?
    • Right-click and choose View Page Source. Is the text you want in that code?
  • The last check sorts pages into two kinds.
  • On a static page, the text you see is in the HTML the site sends. A script can read it directly.
  • On a dynamic page, the browser runs code that fills in the content after the page loads. The page source does not contain that content. Scraping it needs heavier tools.
Live demo

Scraping a practice site

  • Books to Scrape is a site built for practice. It is public and static, and every page has the same layout.
  • One row is one book. The variables are title, price, and star rating.
  • The agent downloads pages 1 and 2, saves their HTML, and extracts the variables into one CSV.
  • To find more pages, click to page 2 and read the address bar. Page 2 is catalogue/page-2.html. Page 3 follows the same pattern.
  • The site has 50 pages. Collect only the pages your question needs.

A site from our projects

  • One volunteer shares a site from their policy area that has no API.
  • Together we run the four checks on it.
  • Is the page static or dynamic? Does it need a login? Does page 2 look like page 1?
  • We open its robots.txt and read what it allows.
  • Could the agent scrape this site? What would you check in what it returns?

What can go wrong

  • Real sites are harder than the practice site:
    • some require a login,
    • some load their content dynamically,
    • some block automated requests,
    • some redesign their pages.
  • After a redesign, the script may return nothing or the wrong part of the page.
  • Compare a few rows with the live page every time you rerun a scraper.
  • Scraping takes longer than you expect. If your project needs it, start now.

Rules for scraping

  • Read the site's robots.txt. Add /robots.txt to the site's address and open it, for example census.gov/robots.txt.
  • Its Disallow lines list pages that automated programs should not visit. The Census Bureau asks programs to stay out of its search results pages.
  • Read the site's terms of use. If they forbid automated collection, use another source.
  • Collect slowly. Pause between requests.
  • Do not collect sensitive personal data.

When a request fails

HTTP status code reference

CodeWhat happenedWhat to do
403the server refused youRead the site's terms. A refusal is often deliberate.
404the address does not existCheck the documentation for the current address.
429too many requests, too fastSlow down and retry later.
500the server itself failedWait and try again.
  • The agent sees the code too. Ask it what the code means for this source.
  • If a source refuses you, write that down. Then look for the same numbers published elsewhere.

Several sources at once · do it on the right

A script loop or subagents?

  • One question decides it. Are the steps the same for every piece?
  • Same steps, such as one API called for 20 agencies: use a loop in a script.
  • Different steps, such as five state websites with different layouts: use one subagent per source, as in Week 3.
  • Give every subagent the same questions:
    • how to collect the data,
    • which fields exist,
    • which years and places are covered,
    • what request limits apply.
  • Compare their reports. Open the documentation links before you choose a source.

A collection you will repeat is a Skill

  • Some collections repeat: this month's new rules, then next month's.
  • Turn the collection into a Skill, as in Week 2. Next time it runs from one command.
  • A classmate can run the same Skill on their own machine.

Project and policy discussion

Your data and your question

  • We go around the room. Each person shows the data for their project to the class.
  • Do you already have the data you need? Or do you need to collect them?
  • After today, would you rethink that? An API or a scraper may let you collect more than you planned.
  • What could more data add to your analysis? What would still be missing?

Assignments

Checkpoint 2: data collection · 5% · individual · due Oct 15 before class. Adapt the classroom collection to your own project’s source. Submit through your GitHub repository: the collection script, the raw data or instructions for getting it, and a README that says how to rerun the collection. Keep the Federal Register abstracts. Week 5 uses them for text coding.

Weekly memo · due Oct 15 before class. Write 1–2 specific observations about using AI to find or collect data for your policy question. Did it help, partly help, or fail to help? Describe a source suggestion, collection problem, or data check, and explain what you learned about whether the data fit your question. Aim for 150–200 words total. Describe what you actually tried and observed in your own words. Do not use AI to write the memo. Begin in class and submit to Canvas Discussions. See assignments and grading.