Week 4: Python Packaging and Documentation with Sphinx#

So far your code has lived in repositories that only you run. A package is code organized so that a teammate can pip install it and build on it, and documentation is what makes that possible. This is the week that turns your coursework into something you can hand a hiring manager, and the paper is Martin and Shi (2025): the deliverable in HW 3 is an installable, tested package that recovers the market’s own probability distribution from option prices.

Announcements#

  • HW 1 is due tonight, at the start of class.

  • HW 2 is due next Tuesday, October 27. Questions at the start of class.

  • HW 3 launches today, due Tuesday, November 3, in week 6: option-implied crash probabilities.

  • Project consultations start this week and run through week 6. If your group has not booked its consultation, do it now.

  • Databento API keys. HW 3 and HW 4 use the key you set up for HW 2. Everything we use is on the CME Globex feed (GLBX.MDP3) and stays inside the course license. The license covers the full order book (mbo, mbp-10) only for roughly the most recent month; older order-book data is billed, so pull a recent date.

Objectives#

  • Explain the anatomy of a published Python package: pyproject.toml, a build backend, versioning, dependencies, and CLI entry points (Writing and Publishing Your Own Python Packages).

  • Build and publish a package with hatch, and know what TestPyPI is for.

  • Explain what semantic versioning promises, and what a breaking change obliges you to do.

  • Generate documentation from the code itself with Sphinx, and say what role docstrings play in it.

  • Contribute to a package you do not own: fork, branch, PR, review, merge.

  • Consume data another repository produces through a ChartBook catalog (The ChartBook Catalog).

  • Know the options data on offer and their quirks: the OptionMetrics volatility surface via WRDS, and CME Globex via Databento.

Agenda Item 1: Anatomy of a Package#

Work through Writing and Publishing Your Own Python Packages: pyproject.toml as the single source of truth, the Hatchling build backend, dynamic versioning, and the one-line [project.scripts] entry that creates a CLI command. Then the packages you can practice on:

  • finm (docs) is the course’s own published package: finm.fixedincome (the yield-curve functions you imported in HW 2), finm.data (Fama-French, Federal Reserve, He-Kelly-Manela and other loaders), and finm.analytics. It is the low-stakes place to practice the contribution loop on a package you do not own, crossing a repository boundary you do not control.

  • stockbeta is a practice package published to TestPyPI and not to the real PyPI. TestPyPI is the rehearsal space: the full publish workflow with none of the consequences of claiming a real name.

  • chartbook is the package you have been installing all quarter. Reading its changelog is a live case study in how a release is actually cut: changelog entry, version bump, hatch build producing the wheel and sdist, publish. Semantic versioning as a contract, and why a breaking change in a 0.x package bumps the minor version.

Agenda Item 2: Documentation with Sphinx#

Sphinx is the tool that builds the documentation for pandas, and for this textbook. It generates browsable reference pages from your code, so the docs sit next to the code and do not fall behind it.

  • Docstrings as the source: what autodoc reads and what conventions make it readable.

  • Forking a project and standing up its docs, featuring jmbejara/quantstats_lumi.

  • Your final project is the natural first candidate for a documented, installable package.

Agenda Item 3: The ChartBook Catalog#

The question a package forces on you is the same one the class report will force in week 6: what should your code do when it needs data another repository already produces? Copying pull code duplicates work and drifts. The ChartBook Catalog is the answer.

  • Pipelines vs. catalogs: a catalog is a chartbook.toml with a [pipelines] registry. The FTSFR organization is the production-scale example: about twenty pipeline repos plus one catalog repo. What the project is, and why every dataset in it rebuilds from source with one command, is in Financial Time Series Forecasting Repository.

  • Hands-on: clone a couple of FTSFR pipeline repos, register them in your global catalog (chartbook catalog add, or members auto-discovery), build and browse the combined site, and load another repo’s dataframe by name with data.load(pipeline=..., dataframe=...).

  • The chapter lists which FTSFR repos run on free public data or WRDS alone. The Bloomberg-gated ones are off the menu.

Agenda Item 4: Options Data and the Options Case Study#

Two data vendors, then the options material the homework builds on.

  • Databento and the hands-on Pulling Market Data From Databento: schemas, symbology, and cost control. CME Globex is free under our license, and it is where the E-mini S&P 500 options in HW 3 come from.

  • LSEG Datastream, briefly, for the vendor landscape.

  • The options case study, finm-32800/case_study_options, reviewed through SPX Hedging, which harmonizes the WRDS OptionMetrics panel with the Databento E-mini chain. This is class material, not the assignment; it establishes the vocabulary the assignment assumes.

Agenda Item 5: Launch HW 3, Option-Implied Crash Probabilities#

HW 3 follows Martin and Shi (2025), Forecasting Crashes with a Smile. You recover the risk-neutral distribution of returns from the volatility smile by the Breeden-Litzenberger method, turn it into an option-implied crash probability, and apply the paper’s corrections to get a bound that forecasts out of sample.

The arc is the lesson. Build the density from the OptionMetrics volatility surface and get a number. Discover that the surface’s observed strikes stop well short of the crash thresholds, so at short horizons the number rests on an extrapolation convention. Rebuild the density from raw quotes on E-mini S&P 500 options from Databento, where every listed strike is quoted, and meet the convexity violations the vendor’s smoothed surface hid from you. Two data sources, one seam between them, and the seam is the point.

The code you complete is a small, pip-installable package with tests, which is why the assignment sits in this week. Smile fitting and density extraction is daily work on an options market-making desk.

Looking ahead to Week 5#

Week 5 is unit tests and data validation: how a pipeline checks its own data before it publishes anything. The opening example is the cleaning of FINRA TRACE, where every filter is a convention, and HW 4 then moves to a case with a ground truth: rebuilding the E-mini order book from the raw message feed and validating it against the vendor’s own book.