Week 2: Task Runners, ChartBook, and SQL, featuring Fama-French 1993#
Last week you pulled data once, by hand. A replication is not one pull: it is pull, clean, merge, construct, test, report, and once a project has more than two steps, running them by hand in the right order becomes the main source of irreproducibility. The fix is a task runner, and the case study that motivates it is the Fama-French (1993) replication that finishes HW 1.
Announcements#
HW 1 is due Tuesday, October 20, but aim to finish it by next week, October 13, when HW 2 launches. We take questions at the start of class. Part D, the Fama-French factors and the investment sort, is what we cover tonight.
Final project list and survey. The Potential Final Projects list is posted, and the preference survey goes out this week. Form your group now: each group is exactly 4 people, and one person per group submits the survey. Assignments are emailed after it closes.
HW 2 launches next week, so HW 1 and HW 2 overlap by a week. Every assignment from here on overlaps the next one that way, which is why the recommended HW 1 finish is a week before its deadline.
Get a Databento API key before October 13. HW 2 pulls 30-Day Fed Funds futures from Databento’s CME Globex feed, which the program’s subscription covers, and the pipeline does not run without a key in your
.env. HW 3 and HW 4 use the same key. See Databento.
Objectives#
State what the CAPM fails to explain and how SMB and HML absorb it: From the CAPM to Multifactor Models. Then construct the factors from raw CRSP and Compustat.
Write enough SQL to query CRSP and Compustat and to make the joins these queries need: Basics of SQL.
Explain what a build system does and why
doitre-runs only what changed: Build Systems and Task Runners.Write
doittasks with correct file dependencies and targets, and read adodo.pyas a dependency graph: PyDoit Examples.Use a notebook for what it is good at, reporting, and keep pipeline logic out of it: Reports with Jupyter Notebooks. Execute a notebook from the command line with
nbconvertand wire it intodoit.Register a project’s notebooks and dataframes in a
chartbook.tomland build a browsable site from it, as one moredoittarget.Know the layout a data project should have, and why, and start one with
chartbook init: the ChartBook project template.Know the alternatives to conda for pinning an environment, and set up a project with one:
uvandpixi.
Agenda#
HW 1 questions. Parts A to C should be working; if your WRDS pull is not authenticating, we fix that first.
The CAPM, in one picture. Last week’s From Mean-Variance to the CAPM has grown since class. Step 2 now splits a stock’s risk into a market part and a firm-specific part, with Ford in 2008 as the example, and the page ends with the four lines that textbooks draw, CAL, CML, SCL, and SML, side by side from ten CRSP stocks. We review the derivation through that figure. The alpha it marks is the intercept of one stock’s history and the distance from the line every stock should sit on, and it is what tonight’s tests look for.
Why the factors exist, before how they are built. The CAPM ended with one factor. If investors also hedge changes in their opportunities, more appear: From the CAPM to Multifactor Models states the ICAPM and shows where Fama and French (1993) fit. Then HW 1 Guide E runs the test on the portfolios your pipeline builds: the market alone is not the tangency portfolio, the CAPM leaves alphas, and SMB and HML absorb some but not all. That is the argument of Fama and French (1993), and the reason we build the factors next.
The case study, end to end. Walk the Fama-French pipeline with HW 1 Guide D: automated CRSP and Compustat pulls feeding book equity, the exchange and share-code screens, NYSE breakpoints, the six size/book-to-market portfolios, and unit tests against the Ken French library. Those steps have to run in order every time the data changes, and that is the problem the rest of tonight solves.
Just enough SQL. The pipeline starts with a pull. In class I show only what is needed to query CRSP and Compustat and to make the simple joins our queries require. Work through Basics of SQL on your own for the rest. → HW 1 Part D: you complete the
WHEREclauses of the Compustat and CRSP-Compustat link queries insrc/pull_CRSP_Compustat.py. These are tested against a small in-memory database, so you can iterate without WRDS.Why task runners? Build Systems and Task Runners: what problem they solve and where
doitsits relative to Make and friends. Then hands-on with PyDoit Examples, using thepydoit/directory of the in-class examples repo. Examples 01 and 02: tasks, file dependencies, targets, and why a correct dependency graph is the whole point. → HW 1 Part D: you completetask_calc_Fama_French_1993andtask_calc_inv_portfoliosindodo.py.Notebooks as tasks. Examples 03 and 04 continue the hands-on. Reports with Jupyter Notebooks: why notebooks are for reporting and not for pipeline steps, and how
nbconvertexecutes one from the command line and exports it to HTML. A notebook is then just anotherdoittask, with file dependencies and targets. → HW 1:task_run_notebooksin yourdodo.pydoes exactly this to the six guide notebooks.ChartBook: the site as a build target. ChartBook catalogs a project’s notebooks, dataframes, and charts and generates a browsable static site from them.
The manifest,
chartbook.toml. Each[notebooks.*],[dataframes.*], and[charts.*]entry says what an output is, where it comes from, and where it lives. Read HW 1’s as the example.The CLI and the Python API.
chartbook buildgenerates the site;chartbook lslists what is registered;chartbook data get-pathreturns a parquet path; anddata.load(pipeline=..., dataframe=...)loads it straight into pandas or polars.Where it sits in the pipeline.
chartbook buildis just anotherdoittask (task_generate_pipeline_sitein HW 1), so the site is regenerated whenever the data changes, not assembled by hand. Next week we publish it to GitHub Pages; in week 4, many projects’ manifests aggregate into one catalog.
Starting a project:
chartbook init. Your final project starts here, so we start one live.The ChartBook project template: the directory layout, and the Cookiecutter Data Science principles behind it (data is immutable, analysis is a DAG, secrets stay out of version control).
chartbook initis a Cookiecutter template. Created with Cruft instead, the project can pull in later improvements to the template withcruft update.The template asks which environment manager to use. That is the cue for
uvandpixi, the follow-up to last week’s conda material: we scaffold one project with each and compare the lockfiles they write.
Final Projects: Preview and Pitch#
This is the week you meet the final projects.
What a finished project looks like. Walk the Project Previews page: three real projects from past cohorts, one each for the site, the report, and the extension, with what to notice and what not to copy.
The full list. Potential Final Projects, organized by topic. Each entry states the exact tables and figures to reproduce and the data sources involved, all verified to be available to you. Skim it before the survey goes out.
Where this can lead. Past final projects from this course grew into the Financial Time Series Forecasting Repository, now a paper with former students as coauthors. We meet it properly in week 4 and fit forecasting models to it in week 9.
Looking ahead to Week 3#
With a pipeline that rebuilds itself, the next question is where its output goes. Week 3 is publishing: the ChartBook site on GitHub Pages and the PDF from LaTeX, plus the pull-request workflow that real teams use to change code. HW 2 launches: the Treasury yield curve and the expected path of the policy rate, published as a website and summarized in a one-page PDF market brief.