# Case Study - FedWatch Replication

Last updated: {sub-ref}`today`


```{toctree}
:maxdepth: 1
:caption: Project Notes

project_overview
```



## Table of Contents

```{toctree}
:maxdepth: 1
:caption: Notebooks 📖
30-Day Fed Funds Futures Data from Databento <cb/notebooks/finm-32800--case_study_fedwatch/01_fed_funds_futures_data>
Replicating the CME FedWatch Tool <cb/notebooks/finm-32800--case_study_fedwatch/02_fedwatch_replication>
```



```{toctree}
:maxdepth: 1
:caption: Pipeline Charts 📈
cb/charts.md
```

```{postlist}
:format: "{title}"
```


```{toctree}
:maxdepth: 1
:caption: Pipeline Dataframes 📊
cb/dataframes/finm-32800--case_study_fedwatch/fed_funds_futures.md
```


```{toctree}
:maxdepth: 1
:caption: Appendix 💡
myst_markdown_demos.md
apidocs/index
```


## Pipeline Specs
| Pipeline Name                   | Case Study - FedWatch Replication                       |
|---------------------------------|--------------------------------------------------------|
| Pipeline ID                     | [finm-32800/case_study_fedwatch](./index.md)              |
| Maintainer                      | Jeremiah Bejarano               |
| Contributors                    | Jeremiah Bejarano |
| Repository                     |                   |
| Pipeline Web Page               | <a href="file:///home/runner/work/case_study_fedwatch/case_study_fedwatch/docs/index.html">Pipeline Web Page      |
| Date of Last Code Update        | 2026-10-11 09:37:36           |
| OS Compatibility                | Windows, Linux, macOS |
| Linked Dataframes               |  [finm-32800/case_study_fedwatch:fed_funds_futures](cb/dataframes/finm-32800--case_study_fedwatch/fed_funds_futures.md)<br>  |


**Build Commands:**
```
doit

```



## About this project

This case study replicates the simplest case of the
[CME FedWatch tool](https://www.cmegroup.com/markets/interest-rates/cme-fedwatch-tool.html):
the market-implied probability of the *next* FOMC rate decision, backed out
of 30-Day Fed Funds futures (ZQ) prices pulled from
[Databento](https://databento.com).

The pipeline (orchestrated with `doit`):

```
pull_fed_funds_futures.py  ->  _data/fed_funds_futures.parquet
                                        |
              +-------------------------+--------------------------+
              |                                                    |
   notebooks 01 & 02 (jupytext)                        fedwatch_chart.py
   executed + rendered to _output/            _output/fedwatch_latest_forecast.{png,html}
```

- `src/fedwatch.py` holds the forecast math as pure, unit-tested functions.
- `src/01_fed_funds_futures_data.ipynb.py` teaches the Databento futures data
  (symbology, schemas, prices as implied rates).
- `src/02_fedwatch_replication.ipynb.py` teaches the FedWatch methodology and
  computes the latest forecast.
- `data_manual/fomc_meetings.csv` is the hand-maintained FOMC calendar —
  **append the new dates each year** when the Fed publishes its schedule
  (see `data_manual/data_README.md`).

## Live site

This repo deploys a live reference copy of the chartbook every morning via
`.github/workflows/deploy_pages.yml` (daily cron at 14:30 UTC, plus every push
to `main`): https://finm-32800.github.io/case_study_fedwatch/

The workflow needs a `DATABENTO_API_KEY` Actions repository secret and
publishes the built `docs/` folder to the `gh-pages` branch; the one-time
setup commands are in a comment at the top of the workflow file.

## Quick Start

First, create a virtual environment and activate it:
```bash
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate
```
Then install the dependencies:
```bash
pip install -r requirements.txt
```
Copy `.env.example` to `.env` and add your Databento API key:
```bash
cp .env.example .env
```
Finally, run the project tasks:
```bash
doit
```
And that's it!

### Data refresh

The data pull is **free**: everything this project downloads is covered by
the course's Databento subscription. As a safety net, the pull script prices
every query with the free `metadata.get_cost` endpoint before downloading
and refuses to run any query whose estimate is not $0.00, so running the
pipeline can never incur a charge.

Once `_data/fed_funds_futures.parquet` exists, `doit` skips the pull. To
refresh with the latest prices and rebuild everything downstream:
```bash
doit forget pull && doit
```

### Other commands

#### Unit Tests and Doc Tests

You can run the unit tests, including doctests, with the following command:
```
pytest --doctest-modules
```

#### Setting Environment Variables

You can [export your environment variables](https://stackoverflow.com/questions/43267413/how-to-set-environment-variables-from-env-file)
from your `.env` files like so, if you wish. This can be done easily in a Linux or Mac terminal with the following command:
```bash
set -a  # automatically export all variables
source .env
set +a
```
On Windows (PowerShell):
```powershell
Get-Content .env | ForEach-Object { if ($_ -match '^([^=]+)=(.*)$') { [Environment]::SetEnvironmentVariable($matches[1], $matches[2], 'Process') } }
```

### Formatting

This project uses [Ruff](https://docs.astral.sh/ruff/) for linting and formatting Python code.

```bash
# Auto-fix linting issues (e.g., unused imports, undefined names)
ruff check . --fix

# Format code (consistent style, spacing, line length)
ruff format .

# Sort imports, then fix linting issues, then format
ruff format . && ruff check --select I --fix . && ruff check --fix .
```

- `ruff check --fix` applies safe auto-fixes for linting violations
- `ruff format` formats code similar to Black
- `--select I` targets only import sorting rules (isort-compatible)

### General Directory Structure

 - The `assets` folder is used for things like hand-drawn figures or other
   pictures that were not generated from code. These things cannot be easily
   recreated if they are deleted. Screenshots used by the notebooks live in
   `src/assets/` (see the README there).

 - The `_output` folder, on the other hand, contains dataframes and figures that are
   generated from code. The entire folder should be able to be deleted, because
   the code can be run again, which would again generate all of the contents.

 - The `data_manual` is for data that cannot be easily recreated. This data
   should be version controlled. Anything in the `_data` folder or in
   the `_output` folder should be able to be recreated by running the code
   and can safely be deleted.

 - I'm using the `doit` Python module as a task runner. It works like `make` and
   the associated `Makefile`s. To rerun the code, install `doit`
   (https://pydoit.org/) and execute the command `doit` from the root
   directory. Note that doit is very flexible and can be used to run code
   commands from the command prompt, thus making it suitable for projects that
   use scripts written in multiple different programming languages.

 - I'm using the `.env` file as a container for absolute paths that are private
   to each collaborator in the project. You can also use it for private
   credentials, if needed. It should not be tracked in Git.

### Data and Output Storage

I'll often use a separate folder for storing data. Any data in the data folder
can be deleted and recreated by rerunning the PyDoit command (the pulls are in
the dodo.py file). Any data that cannot be automatically recreated should be
stored in the "data_manual" folder. Because of the risk of manually-created data
getting changed or lost, I prefer to keep it under version control if I can.
Thus, data in the "_data" folder is excluded from Git (see the .gitignore file),
while the "data_manual" folder is tracked by Git.

Output is stored in the "_output" directory. This includes dataframes, charts, and
rendered notebooks. When the output is small enough, I'll keep this under
version control. I like this because I can keep track of how dataframes change as my
analysis progresses, for example.

Of course, the _data directory and _output directory can be kept elsewhere on the
machine. To make this easy, I always include the ability to customize these
locations by defining the path to these directories in environment variables,
which I intend to be defined in the `.env` file, though they can also simply be
defined on the command line or elsewhere. The `settings.py` is responsible for
loading these environment variables and doing some preprocessing on them.
The `settings.py` file is the entry point for all other scripts to these
definitions. That is, all code that references these variables and others are
loaded by importing `config`.

### Naming Conventions

 - **`pull_` vs `load_`**: Files or functions that pull data from an external
 data source are prepended with "pull_", as in "pull_fed_funds_futures.py".
 Functions that load data that has been cached in the "_data" folder are
 prepended with "load_". For example, `pull_fed_funds_futures.py` contains
 both a `pull_fed_funds_futures` function (hits the Databento API) and a
 `load_fed_funds_futures` function (reads the cached parquet from "_data").

### Dependencies and Virtual Environments

#### Working with `pip` requirements

This project uses `pip` with a virtual environment. Install requirements with:
```bash
pip install -r requirements.txt
```

To update the requirements file after adding new packages:
```bash
pip freeze > requirements.txt
```
