Data Pipelines for Quantitative Research#
By Jeremy Bejarano
Last updated: Oct 6, 2026, 9:22:44 AM CDT
This set of notes is designed to accompany FINM 32800: Data Pipelines for Quantitative Research.
Course Description Data Pipelines for Quantitative Research is a hands-on course centered on building reproducible analytical pipelines: automated and fully reproducible workflows that carry a quantitative analysis from raw data to published result. This course examines every stage of the pipeline, from data extraction and cleaning (extract, transform, load, or ETL), through data validation, exploratory analysis, visualization, modeling, and finally to publication and deployment. In industry, this work is sometimes called productionizing research, pairing developers with researchers to translate ideas into practice. The course teaches the core set of tools used to build such pipelines, tools which are common across computing and data science: build automation and CI/CD, dependency management, SQL, unit testing and automated data-quality checks, the Linux command line, SSH and working with remote machines, Git for version control, peer review through GitHub pull requests, and experiment tracking and model monitoring (basic MLOps).
These skills are taught through a series of case studies, each of which introduces students to a new set of tools and a new key financial data set: pricing and fundamentals from CRSP and Compustat, options data from OptionMetrics, corporate bond transactions from FINRA TRACE, intraday trades and quotes from NYSE TAQ, and order book data from CME Globex.
Prior experience at an intermediate level with Python and the PyData stack is assumed.
Syllabus: web version or PDF download.
Table of Contents#
Lectures 📖
- Lecture 0: Portfolio Selection and a First Look at the Course
- Week 1: Git, GitHub, and Virtual Environments
- Week 2: Task Runners, ChartBook, and SQL, featuring Fama-French 1993
- WARNING: Notes subject to change after this week
- Week 3: Publishing — GitHub Pages, LaTeX, and Pull Requests
- Week 4: Python Packaging and Documentation with Sphinx
- Week 5: Unit Tests and Data Validation with pytest
- Week 6: SQL at Scale, Remote Machines, and Job Schedulers
- Week 7: Orchestration Across Projects with Apache Airflow
- Week 8: CI/CD with GitHub Actions
- Week 9: Exam, and Basic MLOps — Experiment Tracking and Monitoring Models