Case Study - FedWatch Replication#
Last updated: Oct 11, 2026, 2:37:43 PM
Project Notes
Table of Contents#
Pipeline Charts 📈
Pipeline Dataframes 📊
Pipeline Specs#
Pipeline Name |
Case Study - FedWatch Replication |
|---|---|
Pipeline ID |
|
Maintainer |
Jeremiah Bejarano |
Contributors |
Jeremiah Bejarano |
Repository |
|
Pipeline Web Page |
|
Date of Last Code Update |
2026-10-11 09:37:36 |
OS Compatibility |
Windows, Linux, macOS |
Linked Dataframes |
Build Commands:
doit
About this project#
This case study replicates the simplest case of the CME FedWatch tool: the market-implied probability of the next FOMC rate decision, backed out of 30-Day Fed Funds futures (ZQ) prices pulled from Databento.
The pipeline (orchestrated with doit):
pull_fed_funds_futures.py -> _data/fed_funds_futures.parquet
|
+-------------------------+--------------------------+
| |
notebooks 01 & 02 (jupytext) fedwatch_chart.py
executed + rendered to _output/ _output/fedwatch_latest_forecast.{png,html}
src/fedwatch.pyholds the forecast math as pure, unit-tested functions.src/01_fed_funds_futures_data.ipynb.pyteaches the Databento futures data (symbology, schemas, prices as implied rates).src/02_fedwatch_replication.ipynb.pyteaches the FedWatch methodology and computes the latest forecast.data_manual/fomc_meetings.csvis the hand-maintained FOMC calendar — append the new dates each year when the Fed publishes its schedule (seedata_manual/data_README.md).
Live site#
This repo deploys a live reference copy of the chartbook every morning via
.github/workflows/deploy_pages.yml (daily cron at 14:30 UTC, plus every push
to main): https://finm-32800.github.io/case_study_fedwatch/
The workflow needs a DATABENTO_API_KEY Actions repository secret and
publishes the built docs/ folder to the gh-pages branch; the one-time
setup commands are in a comment at the top of the workflow file.
Quick Start#
First, create a virtual environment and activate it:
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
Then install the dependencies:
pip install -r requirements.txt
Copy .env.example to .env and add your Databento API key:
cp .env.example .env
Finally, run the project tasks:
doit
And that’s it!
Data refresh#
The data pull is free: everything this project downloads is covered by
the course’s Databento subscription. As a safety net, the pull script prices
every query with the free metadata.get_cost endpoint before downloading
and refuses to run any query whose estimate is not $0.00, so running the
pipeline can never incur a charge.
Once _data/fed_funds_futures.parquet exists, doit skips the pull. To
refresh with the latest prices and rebuild everything downstream:
doit forget pull && doit
Other commands#
Unit Tests and Doc Tests#
You can run the unit tests, including doctests, with the following command:
pytest --doctest-modules
Setting Environment Variables#
You can export your environment variables
from your .env files like so, if you wish. This can be done easily in a Linux or Mac terminal with the following command:
set -a # automatically export all variables
source .env
set +a
On Windows (PowerShell):
Get-Content .env | ForEach-Object { if ($_ -match '^([^=]+)=(.*)$') { [Environment]::SetEnvironmentVariable($matches[1], $matches[2], 'Process') } }
Formatting#
This project uses Ruff for linting and formatting Python code.
# Auto-fix linting issues (e.g., unused imports, undefined names)
ruff check . --fix
# Format code (consistent style, spacing, line length)
ruff format .
# Sort imports, then fix linting issues, then format
ruff format . && ruff check --select I --fix . && ruff check --fix .
ruff check --fixapplies safe auto-fixes for linting violationsruff formatformats code similar to Black--select Itargets only import sorting rules (isort-compatible)
General Directory Structure#
The
assetsfolder is used for things like hand-drawn figures or other pictures that were not generated from code. These things cannot be easily recreated if they are deleted. Screenshots used by the notebooks live insrc/assets/(see the README there).The
_outputfolder, on the other hand, contains dataframes and figures that are generated from code. The entire folder should be able to be deleted, because the code can be run again, which would again generate all of the contents.The
data_manualis for data that cannot be easily recreated. This data should be version controlled. Anything in the_datafolder or in the_outputfolder should be able to be recreated by running the code and can safely be deleted.I’m using the
doitPython module as a task runner. It works likemakeand the associatedMakefiles. To rerun the code, installdoit(https://pydoit.org/) and execute the commanddoitfrom the root directory. Note that doit is very flexible and can be used to run code commands from the command prompt, thus making it suitable for projects that use scripts written in multiple different programming languages.I’m using the
.envfile as a container for absolute paths that are private to each collaborator in the project. You can also use it for private credentials, if needed. It should not be tracked in Git.
Data and Output Storage#
I’ll often use a separate folder for storing data. Any data in the data folder can be deleted and recreated by rerunning the PyDoit command (the pulls are in the dodo.py file). Any data that cannot be automatically recreated should be stored in the “data_manual” folder. Because of the risk of manually-created data getting changed or lost, I prefer to keep it under version control if I can. Thus, data in the “_data” folder is excluded from Git (see the .gitignore file), while the “data_manual” folder is tracked by Git.
Output is stored in the “_output” directory. This includes dataframes, charts, and rendered notebooks. When the output is small enough, I’ll keep this under version control. I like this because I can keep track of how dataframes change as my analysis progresses, for example.
Of course, the _data directory and _output directory can be kept elsewhere on the
machine. To make this easy, I always include the ability to customize these
locations by defining the path to these directories in environment variables,
which I intend to be defined in the .env file, though they can also simply be
defined on the command line or elsewhere. The settings.py is responsible for
loading these environment variables and doing some preprocessing on them.
The settings.py file is the entry point for all other scripts to these
definitions. That is, all code that references these variables and others are
loaded by importing config.
Naming Conventions#
pull_vsload_: Files or functions that pull data from an external data source are prepended with “pull_”, as in “pull_fed_funds_futures.py”. Functions that load data that has been cached in the “data” folder are prepended with “load”. For example,pull_fed_funds_futures.pycontains both apull_fed_funds_futuresfunction (hits the Databento API) and aload_fed_funds_futuresfunction (reads the cached parquet from “_data”).
Dependencies and Virtual Environments#
Working with pip requirements#
This project uses pip with a virtual environment. Install requirements with:
pip install -r requirements.txt
To update the requirements file after adding new packages:
pip freeze > requirements.txt