Course Map: Is This the Right Course for You?#

Who this course is for#

The program offers three courses that satisfy the computing requirement. They overlap in tools, but they are aimed at different jobs.

  • This course, Data Pipelines for Quantitative Research, is built for students headed toward quantitative research. The work product is an analysis: a replicated result from a finance paper, produced by an automated pipeline that anyone can rerun from raw data to finished report.

  • The other two computing courses are built for students headed toward quantitative development: writing and shipping software for trading and finance.

If you are unsure, ask yourself which question you would rather answer at work. “Does this signal survive out of sample, and can I prove how I computed it?” points here. “How do I make this system fast, correct, and maintainable?” points to the other two. Many of the tools are the same: Git, testing, automation, packaging. The difference is what you build with them.

The course is organized around research papers. Each week we replicate a result from a well-known finance paper, usually with a new data set. The list runs from Markowitz (1952) and the CAPM, through Fama and French (1993) and the Treasury yield curve, to corporate bond factors and intraday measures of liquidity. The full schedule is below.

Why a research course teaches data pipelines#

This is still a computing course. Every student in the program is required to learn a certain amount of computing, and students headed toward research are no exception. Quantitative research is done in code and on data, and it depends on a core set of data science skills: pulling and cleaning data, automating an analysis from end to end, testing it, and publishing a result that someone else can rebuild. A result that cannot be rebuilt from its raw data cannot be trusted, in a journal or on a trading desk.

The ten postings quoted below were picked to illustrate the point. For the wider version, see what firms ask for: a few thousand postings from 64 firms, collected from their own job boards, with the sample frame and its limits written down. Building that dataset is also a worked example of the kind of pipeline this course is about.

These skills are also sought after in industry, and the evidence is in the job postings below. The postings are for research engineers and data engineers, the people who work alongside quantitative researchers. Read them in two ways. If you become a quantitative researcher, these are the skills that let you carry your own idea from raw data to a finished result, and they are the skills of the people you will work with every day. If you end up in a role next to research instead, these are the skills you will be hired for.

Consider this excerpt from Citadel Securities’ posting for a Senior Research Engineer (Data).

“Partner with researchers to produce high-value datasets”

“Translate high-level market research concepts into scalable processes that further transform the data”

This is the partnership that the course is built around. Different companies assign different titles to the engineering side of it. Commonly used titles include research engineer, data engineer, and quantitative developer on a data team, but the work is the same: take research that lives in a researcher’s head or in a published paper and turn it into a production data pipeline that runs reliably every day. Industry’s word for this is productionizing research. In this course you sit on both sides of the partnership. You work through the research in a published paper, and you build the pipeline that reproduces it.

The same skills appear across firms. Consider three recent postings from some of the most selective firms in quantitative finance. The skills they ask for are the skills this course teaches: pipelines that ingest and transform data, SQL and DataFrame libraries, Python on Linux, and automated data-quality checks.

Citadel Securities' Senior Research Engineer (Data) posting, showing the company logo, the responsibilities, and the skills and qualifications

Fig. 11 Citadel Securities, Senior Research Engineer (Data), Miami. (Full archived posting)#

In Citadel’s posting, look for these lines.

  • “Partner with researchers to produce high-value datasets”

  • “Design, create, automate, and maintain custom data pipelines”

  • “Setup ‘data checks’ and alerts to determine when the data is ‘bad’”

  • “Experience with ETL dev”

  • “Strong coding skills: proficiency in Python, SQL DBs, Cloud, schedulers, containers, CI/CD, software packaging”

  • “Data Build Tool (DBT)”

Jane Street's Data Engineer posting, showing the company logo and the About the Position and About You sections

Fig. 12 Jane Street, Data Engineer. (Full archived posting)#

In Jane Street’s posting, look for these lines.

  • “build pipelines that ingest and transform external data”

  • “Proficient with SQL or DataFrame libraries like pandas or Polars”

Hudson River Trading's Data Production Engineer posting, showing the company logo, responsibilities, and qualifications

Fig. 13 Hudson River Trading, Data Production Engineer. (Full archived posting)#

In Hudson River Trading’s posting, look for these lines.

  • “automate tasks using a modern Python data stack”

  • “Perform data reconciliations, validations, and quality checks”

  • “Experience managing ETL pipelines is a plus”

  • “Experienced in at least one SQL dialect (PostgreSQL, MSSQL, MYSQL) and able to use others as needed”

These three are not outliers. The same language appears at many other firms. Postings from ten firms in all are archived on this website, each with its complete text, a screenshot, a PDF snapshot, and a Wayback Machine link. The full list is at the bottom of this page, and the section What the postings ask for, week by week matches their language to the course schedule.

One paper and one tool each week#

Every week pairs a tool for building data pipelines with a well-known finance paper and, usually, a new data set. Replicating the paper gives you repetitions with the tool. It also gives you a working knowledge of the papers and the data that quantitative researchers are expected to know.

Replicating a paper also mimics the task described in the job posting above. The paper plays the role of the researcher. It hands you a high-level research concept, and you translate it into a scalable, automated process that carries raw data to a finished, reproducible result.

The plan for this quarter is below. The second half of the schedule may still change.

Week

Pipeline tool

Paper

Data

0

requirements.txt; clone and run

Markowitz (1952), portfolio selection

CRSP extract

1

Git, GitHub, virtual environments

Sharpe (1964), the CAPM: build the market portfolio and the S&P 500

CRSP

2

Task runners (PyDoit)

Fama and French (1993), the three-factor model

CRSP, Compustat

3

Reproducible reports: notebooks, GitHub Pages, LaTeX

Gürkaynak, Sack, and Wright (2006), the Treasury yield curve, together with the expected path of the policy rate read from fed funds futures (the CME FedWatch method)

CRSP Treasuries, CME fed funds futures, FRED

4

Python packaging and documentation

Martin and Shi (2025), forecasting crashes with a smile: option-implied crash probabilities from the volatility surface

OptionMetrics, CME Globex

5

Unit tests and data validation

Easley, López de Prado, and O’Hara (2012), flow toxicity and liquidity: rebuilding the order book and measuring toxicity in E-mini futures

CME Globex order book (Databento)

6

SQL at scale, remote machines, and job schedulers: the WRDS Cloud and Midway

Holden and Jacobsen (2014), measuring liquidity from intraday trades and quotes

NYSE TAQ

7

Orchestration across projects: Apache Airflow

Goyal, Welch, and Zafirov (2024), the equity premium predictors

FRED and other public data; the FTSFR pipelines

8

CI/CD with GitHub Actions

Bernanke and Kuttner (2005), monetary policy surprises from fed funds futures

CME fed funds futures

9

Basic MLOps: experiment tracking and monitoring models

Bejarano et al. (2026), an open benchmark for forecasting across financial markets

The FTSFR datasets

The arc on the finance side runs from equities (weeks 0 to 2), to interest rates and options (weeks 3 and 4), to market microstructure (weeks 5 and 6), and ends with return prediction (weeks 7 to 9), which draws on everything before it. Corporate bonds and the cleaning of FINRA TRACE appear in week 5 as the opening example of data validation.

The final project#

The final project is a replication of a published paper, done in a small group, and built as a complete pipeline: one command takes it from raw data to a finished report. Browse the past final projects to see what students have built.

What the postings ask for, week by week#

Each week’s tool was chosen because firms ask for it by name. Below, each week of the schedule above is matched to what the postings say, quoted verbatim. A single sentence in a posting often spans several of the course’s topics, so some quotes appear under more than one week. Each firm name links to the full archived posting.

Week 1: Git, GitHub, and virtual environments

  • “Demonstrated experience working on an Agile team employing software engineering best practices, such as GitOps and CI/CD, to deliver complex software projects” (Akuna Capital)

  • “Strong command of engineering best practices, including code quality, design documentation, peer code reviews, automated testing, and test coverage” (AQR)

Week 2: Task runners (PyDoit) and ETL

  • “Design, create, automate, and maintain custom data pipelines” (Citadel Securities)

  • “build pipelines that ingest and transform external data” (Jane Street)

  • “3+ years of demonstrated experience designing and implementing ingestion pipelines” (DRW)

  • “Improve data ETL pipeline and build tools to analyze new data efficiently.” (Point72)

Week 3: Reproducible reports

  • “Setup analytical infrastructure to facilitate exploration and visualization of datasets” (Citadel Securities)

  • “Clear written and verbal communication skills to translate complex technical work for business stakeholders and collaborate in agile, cross functional teams.” (JPMorganChase)

  • “Strong ability to communicate with other stakeholders (e.g., data vendors, QRs, etc.)” (Citadel Securities)

Week 4: Python packaging and documentation

  • “Strong coding skills: proficiency in Python, SQL DBs, Cloud, schedulers, containers, CI/CD, software packaging” (Citadel Securities)

  • “gain experience with our full-cycle process for development, testing, and release” (Jump Trading)

  • “Produce clean, well-tested, and documented code with a clear design to support mission critical applications” (Akuna Capital)

Week 5: Unit tests and data validation

  • “Setup ‘data checks’ and alerts to determine when the data is ‘bad’” (Citadel Securities)

  • “Proven expertise in developing data quality control processes to detect gaps or inaccuracies” (DRW)

  • “Perform data reconciliations, validations, and quality checks” (Hudson River Trading)

  • “Build automated data validation test suites that ensure that data is processed and published in accordance with well-defined Service Level Agreements (SLA’s) pertaining to data quality, data availability and data correctness” (Akuna Capital)

Week 6: SQL at scale, remote machines, and job schedulers

  • “Proficient with SQL or DataFrame libraries like pandas or Polars” (Jane Street)

  • “Experienced in at least one SQL dialect (PostgreSQL, MSSQL, MYSQL) and able to use others as needed” (Hudson River Trading)

  • “Comfortable with the Linux command line” (Hudson River Trading)

  • “Unix scripting experience (bash, python, etc.)” (IMC Trading)

  • “Data Build Tool (DBT)” (Citadel Securities)

Week 7: Orchestration across projects (Apache Airflow)

  • “Strong coding skills: proficiency in Python, SQL DBs, Cloud, schedulers, containers, CI/CD, software packaging” (Citadel Securities)

  • “Translate high-level market research concepts into scalable processes that further transform the data” (Citadel Securities)

Week 8: CI/CD with GitHub Actions

  • “Demonstrated experience working on an Agile team employing software engineering best practices, such as GitOps and CI/CD, to deliver complex software projects” (Akuna Capital)

  • “Ability to work with a team in a fast-paced environment, deploying new software daily” (Jump Trading)

  • “Build, deploy, and monitor our data processing pipelines (Java, Python, Spark, Flink)” (IMC Trading)

Week 9: Basic MLOps: experiment tracking and monitoring models

  • “they engineer data, build and deploy models, and monitor them in production” (JPMorganChase)

  • “Experience with monitoring, observability, and alerting systems for data pipelines” (DRW)

Every week: know the data

  • “An understanding of the financial data vendor landscape and product offerings” (DRW)

  • “Experience with financial datasets (e.g. Refinitiv, S&P, Bloomberg) is a big plus” (Hudson River Trading)

  • “Strong understanding of financial point-in-time and time-series data and analysis” (DRW)

The course does not cover everything in these postings. Kafka, Spark, Kubernetes, and warehouse platforms like Snowflake and Databricks appear in several and are beyond our scope. What the course does claim is the layer beneath all of it: the habits of building pipelines that are versioned, tested, documented, automated, and reproducible.

Appendix: the archived job postings#

Job postings are ephemeral. Once a role is filled the page is taken down. Each posting quoted on this page is therefore preserved on its own page of this website, with the full text, a screenshot, a PDF snapshot, and a link to a Wayback Machine copy. All ten were collected in August 2026.

Firm

Role

Location

Archived copy

Jane Street

Data Engineer

London

Full posting

Citadel Securities

Senior Research Engineer (Data)

Miami

Full posting

IMC Trading

Data Engineer

Chicago

Full posting

DRW

Data Engineer, Cumberland/FICCO

Chicago

Full posting

Point72, Cubist Systematic Strategies

Data Engineer

New York

Full posting

Jump Trading

Campus Data Engineer (Intern)

Chicago

Full posting

Hudson River Trading

Data Production Engineer

London, New York, Singapore

Full posting

Akuna Capital

Software Engineer, Data Engineering

Chicago

Full posting

AQR Capital Management

Portfolio Analytics Engineer, Vice President

Greenwich, CT

Full posting

JPMorganChase

Data & AI Internship Program

Full posting

For the longer argument these postings support, see What Is This Course About? The index of archived postings is here.

References#

The papers in the table above, in the order that we cover them. Each title links to the published version. Where a free version exists, it is linked as well. From a campus network or the library proxy, the journal links should give you the full text.