Skip to main content
Back to top
Ctrl
+
K
Data Pipelines for Quantitative Research
Course Syllabus: FINM 32800, Autumn 2026
Archived Job Postings
Jane Street — Data Engineer
Citadel Securities — Senior Research Engineer (Data)
IMC Trading — Data Engineer
DRW — Data Engineer, Cumberland/FICCO
Point72 (Cubist Systematic Strategies) — Data Engineer
Jump Trading — Campus Data Engineer (Intern)
Hudson River Trading — Data Production Engineer
Akuna Capital — Software Engineer, Data Engineering
AQR Capital Management — Portfolio Analytics Engineer
JPMorganChase — Data & AI Internship Program
Acknowledgments
Appendix
The Bloomberg Terminal
Web Authentication and Authorization
Job Postings: Sample, Methods and Limitations
Lectures 📖
Lecture 0: Portfolio Selection and a First Look at the Course
Course Map: Is This the Right Course for You?
What Firms Ask For
A Tour of the Course
Getting Set Up
Clone and Run: What
requirements.txt
Is For
Portfolio Selection: Markowitz (1952)
Appendix: Deriving the Mean-Variance Frontier
A Tour of the Data
Week 1: Git, GitHub, and Virtual Environments
What Is This Course About? Why Data Pipelines?
What are Reproducible Analytical Pipelines?
Case Study: Is There A Reproducibility Crisis In Finance?
Virtual Environments
Introduction to WRDS
Example: Connecting to the WRDS Platform With Python
Env Files, Secrets, and the Separations of Settings from Code
From Mean-Variance to the CAPM
Week 2: Task Runners, ChartBook, and SQL, featuring Fama-French 1993
From the CAPM to Multifactor Models
Basics of SQL
What is a build system or task runner?
PyDoit Examples Walkthrough
Reports with Jupyter Notebooks
Project Structure: “Chartbook” Template
Modern Environment Tools: uv and pixi
WARNING: Notes subject to change after this week
Week 3: Publishing — GitHub Pages, LaTeX, and Pull Requests
GitHub Pages and Tearsheets
Introduction to LaTeX
LaTeX Essentials
GitHub Issues and Pull Requests: Enhancing Collaborative Development
Week 4: Python Packaging and Documentation with Sphinx
Writing and Publishing Your Own Python Packages
Sphinx
The ChartBook Catalog: Data Dependencies Across Repositories
Financial Time Series Forecasting Repository (FTSFR)
Databento
Example: Pulling Market Data From Databento
LSEG Datastream
Case Study: Hedging A Long-Only SPX Portfolio With Costless Collars
Week 5: Unit Tests and Data Validation with pytest
Unit Tests
Data Sources Overview
Cleaning TRACE Corporate Bond Data: A Walkthrough
Week 6: SQL at Scale, Remote Machines, and Job Schedulers
Strategies for Medium-Sized Data
Polars Exercises: Code Snippets
Remote Machines and High-Performance Computing
Exercise: Jupyter Notebook on Midway via SSH
Optional Lab: Clean TRACE on Midway
The Collaborative Report and Style Guides
Week 7: Orchestration Across Projects with Apache Airflow
Week 8: CI/CD with GitHub Actions
Creating a Live Dashboard Example with GitHub Actions
Scheduling Recurring Tasks with Cron
Week 9: Exam, and Basic MLOps — Experiment Tracking and Monitoring Models
Homework 📝
Homework 0
Homework 1: CRSP, the CAPM, and Fama-French
HW 1 Guide A: GitHub Skills
HW 1 Guide B: The CRSP Market Index
HW 1 Guide C: Reconstructing the S&P 500
HW 1 Guide D: Constructing the Fama-French Factors
HW 1 Guide E: Testing the CAPM and the Three-Factor Model
Homework 2: The Yield Curve and the Policy Path
CRSP Treasury Data Oveview
Replicating the Gürkaynak, Sack, and Wright (2006) Treasury Yield Curve
01. 30-Day Fed Funds Futures (ZQ) Data from Databento
02. Replicating the CME FedWatch Tool
Homework 3: Option-Implied Crash Probabilities
Homework 4: Order Book Validation and Flow Toxicity
Homework 5: The Class Report
Final Project
Final Project Instructions and Rubric
Project Previews: What a Finished Project Looks Like
Past Final Projects
List of Potential Final Projects
Exam Preparation
Repository
Open issue
Index