Case Study - FedWatch Replication#

Last updated: Oct 11, 2026, 2:37:43 PM

Project Notes

Table of Contents#

Pipeline Charts 📈

Pipeline Specs#

Pipeline Name

Case Study - FedWatch Replication

Pipeline ID

finm-32800/case_study_fedwatch

Maintainer

Jeremiah Bejarano

Contributors

Jeremiah Bejarano

Repository

Pipeline Web Page

Pipeline Web Page

Date of Last Code Update

2026-10-11 09:37:36

OS Compatibility

Windows, Linux, macOS

Linked Dataframes

finm-32800/case_study_fedwatch:fed_funds_futures

Build Commands:

doit

About this project#

This case study replicates the simplest case of the CME FedWatch tool: the market-implied probability of the next FOMC rate decision, backed out of 30-Day Fed Funds futures (ZQ) prices pulled from Databento.

The pipeline (orchestrated with doit):

pull_fed_funds_futures.py  ->  _data/fed_funds_futures.parquet
                                        |
              +-------------------------+--------------------------+
              |                                                    |
   notebooks 01 & 02 (jupytext)                        fedwatch_chart.py
   executed + rendered to _output/            _output/fedwatch_latest_forecast.{png,html}
  • src/fedwatch.py holds the forecast math as pure, unit-tested functions.

  • src/01_fed_funds_futures_data.ipynb.py teaches the Databento futures data (symbology, schemas, prices as implied rates).

  • src/02_fedwatch_replication.ipynb.py teaches the FedWatch methodology and computes the latest forecast.

  • data_manual/fomc_meetings.csv is the hand-maintained FOMC calendar — append the new dates each year when the Fed publishes its schedule (see data_manual/data_README.md).

Live site#

This repo deploys a live reference copy of the chartbook every morning via .github/workflows/deploy_pages.yml (daily cron at 14:30 UTC, plus every push to main): https://finm-32800.github.io/case_study_fedwatch/

The workflow needs a DATABENTO_API_KEY Actions repository secret and publishes the built docs/ folder to the gh-pages branch; the one-time setup commands are in a comment at the top of the workflow file.

Quick Start#

First, create a virtual environment and activate it:

python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

Then install the dependencies:

pip install -r requirements.txt

Copy .env.example to .env and add your Databento API key:

cp .env.example .env

Finally, run the project tasks:

doit

And that’s it!

Data refresh#

The data pull is free: everything this project downloads is covered by the course’s Databento subscription. As a safety net, the pull script prices every query with the free metadata.get_cost endpoint before downloading and refuses to run any query whose estimate is not $0.00, so running the pipeline can never incur a charge.

Once _data/fed_funds_futures.parquet exists, doit skips the pull. To refresh with the latest prices and rebuild everything downstream:

doit forget pull && doit

Other commands#

Unit Tests and Doc Tests#

You can run the unit tests, including doctests, with the following command:

pytest --doctest-modules

Setting Environment Variables#

You can export your environment variables from your .env files like so, if you wish. This can be done easily in a Linux or Mac terminal with the following command:

set -a  # automatically export all variables
source .env
set +a

On Windows (PowerShell):

Get-Content .env | ForEach-Object { if ($_ -match '^([^=]+)=(.*)$') { [Environment]::SetEnvironmentVariable($matches[1], $matches[2], 'Process') } }

Formatting#

This project uses Ruff for linting and formatting Python code.

# Auto-fix linting issues (e.g., unused imports, undefined names)
ruff check . --fix

# Format code (consistent style, spacing, line length)
ruff format .

# Sort imports, then fix linting issues, then format
ruff format . && ruff check --select I --fix . && ruff check --fix .
  • ruff check --fix applies safe auto-fixes for linting violations

  • ruff format formats code similar to Black

  • --select I targets only import sorting rules (isort-compatible)

General Directory Structure#

  • The assets folder is used for things like hand-drawn figures or other pictures that were not generated from code. These things cannot be easily recreated if they are deleted. Screenshots used by the notebooks live in src/assets/ (see the README there).

  • The _output folder, on the other hand, contains dataframes and figures that are generated from code. The entire folder should be able to be deleted, because the code can be run again, which would again generate all of the contents.

  • The data_manual is for data that cannot be easily recreated. This data should be version controlled. Anything in the _data folder or in the _output folder should be able to be recreated by running the code and can safely be deleted.

  • I’m using the doit Python module as a task runner. It works like make and the associated Makefiles. To rerun the code, install doit (https://pydoit.org/) and execute the command doit from the root directory. Note that doit is very flexible and can be used to run code commands from the command prompt, thus making it suitable for projects that use scripts written in multiple different programming languages.

  • I’m using the .env file as a container for absolute paths that are private to each collaborator in the project. You can also use it for private credentials, if needed. It should not be tracked in Git.

Data and Output Storage#

I’ll often use a separate folder for storing data. Any data in the data folder can be deleted and recreated by rerunning the PyDoit command (the pulls are in the dodo.py file). Any data that cannot be automatically recreated should be stored in the “data_manual” folder. Because of the risk of manually-created data getting changed or lost, I prefer to keep it under version control if I can. Thus, data in the “_data” folder is excluded from Git (see the .gitignore file), while the “data_manual” folder is tracked by Git.

Output is stored in the “_output” directory. This includes dataframes, charts, and rendered notebooks. When the output is small enough, I’ll keep this under version control. I like this because I can keep track of how dataframes change as my analysis progresses, for example.

Of course, the _data directory and _output directory can be kept elsewhere on the machine. To make this easy, I always include the ability to customize these locations by defining the path to these directories in environment variables, which I intend to be defined in the .env file, though they can also simply be defined on the command line or elsewhere. The settings.py is responsible for loading these environment variables and doing some preprocessing on them. The settings.py file is the entry point for all other scripts to these definitions. That is, all code that references these variables and others are loaded by importing config.

Naming Conventions#

  • pull_ vs load_: Files or functions that pull data from an external data source are prepended with “pull_”, as in “pull_fed_funds_futures.py”. Functions that load data that has been cached in the “data” folder are prepended with “load”. For example, pull_fed_funds_futures.py contains both a pull_fed_funds_futures function (hits the Databento API) and a load_fed_funds_futures function (reads the cached parquet from “_data”).

Dependencies and Virtual Environments#

Working with pip requirements#

This project uses pip with a virtual environment. Install requirements with:

pip install -r requirements.txt

To update the requirements file after adding new packages:

pip freeze > requirements.txt