Week 4 · Oct 1, 2026 · Thursday, 9:00–11:30 am · 280 Brook Street, Room 110
AI for data collection
Collect a documented dataset and explain how its coverage supports or changes the project question.
Opening
Today
- Why a policy analyst would collect data.
- APIs, with the Federal Register as the example.
- Sources without an API: web scraping.
- Collecting from several sources at once.
- What your data leave out.
Why collect data at all
- Policy analysts usually work from tables that someone else published.
- Last week's reading said the same. Most policy research reuses what others produced.
- A published table may not fit your question. It may have the wrong years, the wrong places, or too little detail.
- With an agent, you can collect the documents your question needs in one class session.
- So the question changes. Are these data worth collecting, and what do they leave out?
Readings
Datasheets for Datasets
- Gebru et al., “Datasheets for Datasets” (2021). Read for what to write down about the data you collect.
APIs · do it on the right
What an API is
- A government website is made for a person who reads one page at a time. A policy analysis often needs hundreds of documents at once.
- An agency can also offer a second way to get the same documents, made for programs. This is an API, short for application programming interface.
- A program sends the API a request. Parameters in the request say which documents you want, such as the agency, the dates, and the document type.
- The documents come back with no page layout, in a form a script can read directly. Every document has the same fields, such as title and date.
- Some APIs require a key. A key identifies you and limits how fast you may send requests.
- Keep your key out of Git. The Federal Register API needs no key.
- Many federal agencies offer an API. Many state agencies and city governments do not. Sources without an API covers what to do then.
The Federal Register
- The Federal Register is the U.S. government's daily journal.
- Every federal agency publishes its proposed rules, final rules, and notices there.
- A proposed rule is a draft that the public can comment on. A final rule is the version that takes effect.
- A proposed rule usually lists a docket number and the date public comments close.
- A policy question it can answer: what did the Environmental Protection Agency propose in 2025, and when could the public respond?
- In 2025 the EPA published 242 proposed rules. On the website, that is 242 pages to open one at a time.
Start on the website
- Open federalregister.gov. Search for proposed rules from the EPA in 2025.
- Open one document. Find the title, the agency, the date, the summary, the docket number, and the comment deadline.
- The API returns those same pieces for every matching document, with no page layout.
- The website is for people. When a script asks the website for a page, it is sent to a block page.
- The API answers the same script.
The API documentation
- The Federal Register API documentation lists what you can ask for and the name of every parameter.
- Agency names must be spelled the way the API spells them.
- In the URL, each agency is written in lowercase with hyphens. The EPA is
environmental-protection-agency. The Department of Education iseducation-department. - The agency list gives this form for every agency, in the field labeled
slug. - Read the documentation before you ask the agent for anything. You need it to check the URL the agent builds.
Two example requests
- Proposed rules from the Environmental Protection Agency, published in 2025:
- Final rules from the Department of Education, published in 2024:
- Everything before the question mark is the endpoint, the address of the API. It is the same in both.
- Everything after it is the parameters. The two URLs differ in three of them: the agency, the year, and the type.
| Parameter | What it sets | Values you can use |
|---|---|---|
conditions[agencies][] | the agency | the agency, as written in the agency list: environmental-protection-agency, education-department |
conditions[publication_date][year] | the year of publication | a four-digit year, such as 2024 |
conditions[type][] | the kind of document | PRORULE for a proposed rule, RULE for a final rule, NOTICE for a notice, PRESDOCU for a presidential document |
per_page | how many documents one page shows | a number, such as 5 |
- Copy the agency from the list. A guess such as
department-of-educationreturns an error. - What comes back is in JSON format. It is built like a Python dictionary: each field name paired with its value.
- Raw JSON is hard to read in a browser. To see its structure, paste it into a formatter such as jsonformatter.org.
Your own example
- Choose an agency that works on your policy topic, such as the Labor Department.
- Open the agency list. Search the page for "Labor". In that agency's entry, the field labeled
slugreadslabor-department. - Start from either example URL above. Put
labor-departmentin the agency parameter. - Pick a year. Pick a type from the table above:
PRORULE,RULE, orNOTICE. - Reload. Run the same search on the website. The two counts should match.
- You can also describe what you want and ask the agent to build the URL. Read every parameter before you run it.
What you do when the agent writes the request
- In a programming course, you would read the API documentation and write the code yourself.
- Here you say what you want in plain words. The agent reads the documentation, writes the code, and runs it.
- Your job is checking:
- read the URL the agent built,
- compare a few documents with the same pages on federalregister.gov,
- check that the script stays under the source's request limit.
- This check matters. On a replication benchmark, agents scored about 96 on running an analysis.
- The same agents scored about 31 on finding the data the analysis needed.
Nguyen et al., "ReplicatorBench" (2026), arXiv:2602.11354.
Every API works the same way
| API | You send | You get back |
|---|---|---|
| Federal Register | an agency, a year, a document type | the matching documents |
| Regulations.gov | a docket number | the public comments on that proposed rule |
| a language model, such as Claude or GPT | a prompt | the model's answer |
- You send a request with parameters. Data come back.
- In Week 5, you send the documents you collect today to the third row.
Keep the raw data
- Save the response exactly as it came back, in
data/. Never edit that file. - Next to it, write down:
- the full request URL,
- the parameters,
- the date and time you collected it.
- Do all cleaning in a script. Then the clean table can always be rebuilt from the raw file.
- This record of where the data came from is called provenance.
Collecting Federal Register documents
- Collect up to 100 rules and proposed rules that one agency published in 2025.
- Note how the API chose those 100: the sort order and the filters.
- Make a table of the documents by month. Label it as the first 100, since the agency may have published more.
- Change the agency and run it again.
- Open a few documents on federalregister.gov. Compare the saved title, date, and abstract with the page.
Public sources with an API
| Source | What is in it |
|---|---|
| Federal Register | rules, proposed rules, and notices from federal agencies, with abstracts and links |
| Bureau of Labor Statistics | employment, wages, and prices, by area and industry |
| Socrata | the platform many city governments use to publish permits, inspections, and service requests |
| data.gov | the federal catalog of datasets, for finding an agency's data |
| USAspending | federal contracts, grants, and other awards |
| EPA and HUD | environmental monitoring and housing program data |
Does your source have an API?
- Before collecting from any website, check how it publishes its data. Check in this order:
- a ready-made download, such as a CSV file,
- an API,
- the web page itself.
- A download is the easiest route. The page itself is the last resort.
- To find an API, search the web for the site's name and "API". Many sites also have a page called "Developers" or "Data".
- If you find an API, open its documentation. Check what you can ask for and whether you need a key.
Most sources have no API
- The Federal Register is the easy case. It has an API and needs no key.
- Many state agencies and city governments offer neither an API nor a download. Their data are only on their web pages.
- Your project's source may be one of these.
- The agent can still collect from those pages. The next section shows how.
Sources without an API · do it on the right
Reading a page directly
- The simplest way is to ask the agent to read the page directly.
- In Claude Code, the tool is called WebFetch.
- In Codex, it is web search. Start Codex with
codex --searchto turn it on.
- You give it a URL and say what you want, such as "get the monthly unemployment table from this page and save it as a CSV."
- What the agent gets back can be incomplete. In Claude Code, a second, smaller model reads the page and passes on only its answer. Long pages are cut off first.
- Every page costs model tokens. In a test for this course, Codex used about 10,000 tokens to read one page.
- Nothing is saved to rerun. There is no script and no copy of the page.
- Use it to look at one page or a few.
- For more pages, or for a collection you will repeat, the agent needs to write a script. That is web scraping.
What web scraping is
- Every web page is built from HTML, the code your browser turns into the page you see.
- Web scraping is a script that downloads a page's HTML and pulls out the pieces you want.
- The pieces can be a table, a list of reports, dates, prices, or the text of each page.
- People scrape when:
- the site has no API,
- its API leaves out what they need,
- the site shows content to readers and offers no way to reuse it.
- The agent writes the script. Your job is to judge whether a page can be scraped, and to check what comes back.
Is a page easy to scrape?
- Run four checks in your browser:
- Can you see the content without logging in?
- Does the content stay the same when you refresh?
- Does page 2 look like page 1?
- Right-click and choose View Page Source. Is the text you want in that code?
- The last check sorts pages into two kinds.
- On a static page, the text you see is in the HTML the site sends. A script can read it directly.
- On a dynamic page, the browser runs code that fills in the content after the page loads. The page source does not contain that content. Scraping it needs heavier tools.
Scraping a practice site
- Books to Scrape is a site built for practice. It is public and static, and every page has the same layout.
- One row is one book. The variables are title, price, and star rating.
- The agent downloads pages 1 and 2, saves their HTML, and extracts the variables into one CSV.
- To find more pages, click to page 2 and read the address bar. Page 2 is
catalogue/page-2.html. Page 3 follows the same pattern. - The site has 50 pages. Collect only the pages your question needs.
A site from our projects
- One volunteer shares a site from their policy area that has no API.
- Together we run the four checks on it.
- Is the page static or dynamic? Does it need a login? Does page 2 look like page 1?
- We open its robots.txt and read what it allows.
- Could the agent scrape this site? What would you check in what it returns?
What can go wrong
- Real sites are harder than the practice site:
- some require a login,
- some load their content dynamically,
- some block automated requests,
- some redesign their pages.
- After a redesign, the script may return nothing or the wrong part of the page.
- Compare a few rows with the live page every time you rerun a scraper.
- Scraping takes longer than you expect. If your project needs it, start now.
Rules for scraping
- Read the site's robots.txt. Add
/robots.txtto the site's address and open it, for example census.gov/robots.txt. - Its
Disallowlines list pages that automated programs should not visit. The Census Bureau asks programs to stay out of its search results pages. - Read the site's terms of use. If they forbid automated collection, use another source.
- Collect slowly. Pause between requests.
- Do not collect sensitive personal data.
When a request fails
| Code | What happened | What to do |
|---|---|---|
403 | the server refused you | Read the site's terms. A refusal is often deliberate. |
404 | the address does not exist | Check the documentation for the current address. |
429 | too many requests, too fast | Slow down and retry later. |
500 | the server itself failed | Wait and try again. |
- The agent sees the code too. Ask it what the code means for this source.
- If a source refuses you, write that down. Then look for the same numbers published elsewhere.
Several sources at once · do it on the right
A script loop or subagents?
- One question decides it. Are the steps the same for every piece?
- Same steps, such as one API called for 20 agencies: use a loop in a script.
- Different steps, such as five state websites with different layouts: use one subagent per source, as in Week 3.
- Give every subagent the same questions:
- how to collect the data,
- which fields exist,
- which years and places are covered,
- what request limits apply.
- Compare their reports. Open the documentation links before you choose a source.
A collection you will repeat is a Skill
- Some collections repeat: this month's new rules, then next month's.
- Turn the collection into a Skill, as in Week 2. Next time it runs from one command.
- A classmate can run the same Skill on their own machine.
Project and policy discussion
Your data and your question
- We go around the room. Each person shows the data for their project to the class.
- Do you already have the data you need? Or do you need to collect them?
- After today, would you rethink that? An API or a scraper may let you collect more than you planned.
- What could more data add to your analysis? What would still be missing?
Assignments
Checkpoint 2: data collection · 5% · individual · due Oct 15 before class. Adapt the classroom collection to your own project’s source. Submit through your GitHub repository: the collection script, the raw data or instructions for getting it, and a README that says how to rerun the collection. Keep the Federal Register abstracts. Week 5 uses them for text coding.
Weekly memo · due Oct 15 before class. Write 1–2 specific observations about using AI to find or collect data for your policy question. Did it help, partly help, or fail to help? Describe a source suggestion, collection problem, or data check, and explain what you learned about whether the data fit your question. Aim for 150–200 words total. Describe what you actually tried and observed in your own words. Do not use AI to write the memo. Begin in class and submit to Canvas Discussions. See assignments and grading.