What Firms Ask For#

The course map quotes ten job postings and matches their language to the course schedule. This page widens that from ten postings to a few thousand.

Every few months a script collects job postings from the applicant-tracking systems that quantitative finance firms use to publish their own openings, and counts which technologies those postings name. It is a small, reproducible measurement over a defined set of firms, and it is worth having for two reasons: the patterns in it do not depend on which ten postings someone happened to save, and building it is itself an example of the kind of pipeline this course is about — pull from an API, normalise, extract features, report, and keep the whole thing rebuildable.

This page is deliberately short. A companion methods notebook in the project repository carries the sample frame, the ways of counting, the extraction rules and the limitations, and every number here traces back to a table there.

from jobpostings import report

d = report.load()
report.headline(d)
collection runs: 2026-09-25, 2026-09-27
11,346 postings from 64 firms
9,237 technical postings
1,600 mention a pipeline, ETL/ELT or orchestration

Scheduling and orchestration#

The layer this course spends the most time on. Read the Firms column first: it counts how many distinct firms have at least one posting naming the tool, which is the hardest number to distort. A firm that advertises the same role in six offices, or repeats its stack description across unrelated postings, inflates the posting count and leaves the firm count alone.

Tools named by nobody are listed underneath, because an absence is a finding too.

sched = report.tools_table(d["summary"], layers=report.SCHEDULING_LAYERS)
sched[["Tool", "Layer", "Postings", "Firms", "% firms naming it", "% postings (shrunk)"]]
Tool Layer Postings Firms % firms naming it % postings (shrunk)
0 Apache Airflow orchestrator 164 33 57.9 13.1
1 Dagster orchestrator 23 14 24.6 1.8
2 Slurm hpc_scheduler 26 12 21.1 2.2
3 Argo Workflows / Argo CD durable_workflow 23 10 17.5 1.5
4 CA AutoSys enterprise_scheduler 28 9 15.8 1.7
5 Prefect orchestrator 13 9 15.8 1.1
6 AWS Step Functions durable_workflow 42 8 14.0 2.3
7 BMC Control-M enterprise_scheduler 25 8 14.0 1.8
8 Azure Data Factory durable_workflow 10 6 10.5 0.6
9 Temporal.io durable_workflow 20 5 8.8 1.1
10 Hadoop YARN hpc_scheduler 9 5 8.8 0.6
11 Kueue hpc_scheduler 2 2 3.5 0.2
12 Make / Makefiles build_runner 2 2 3.5 0.1
13 cron / systemd timers enterprise_scheduler 1 1 1.8 0.1
14 HTCondor hpc_scheduler 1 1 1.8 0.1
15 Luigi orchestrator 1 1 1.8 0.1
16 Apache NiFi durable_workflow 1 1 1.8 0.1
17 Tidal Workload Automation enterprise_scheduler 1 1 1.8 0.1
absent = report.tools_table(
    d["summary"], layers=report.SCHEDULING_LAYERS, drop_empty=False
)
print("Scheduling tools named by no firm in the sample:")
for t in sorted(absent[absent["Postings"] == 0]["Tool"]):
    print(f"  {t}")
Scheduling tools named by no firm in the sample:
  Flyte
  GNU Parallel
  IBM Spectrum LSF
  Kedro
  Kestra
  Mage
  Metaflow
  Nextflow
  PBS / Torque
  Snakemake
  pydoit

Airflow is named by more firms than every other orchestrator combined. Luigi, which appears in nearly every “top data pipeline tools” article, is named once in the entire corpus.

The two columns disagree in a way worth understanding. A tool can be named by many firms but in very few of their postings — breadth without depth. Dagster is the clearest case: second only to Airflow by firms, but named in barely more than one posting per firm, almost always as an acceptable alternative inside a list that begins with Airflow, rather than as the thing they build on.

fig, top = report.barh(
    sched,
    ["% firms naming it", "% postings (shrunk)"],
    ["Share of firms naming the tool", "Share of postings (shrunk)"],
    "Scheduling tools named in quant-finance job postings",
)
lead, runner = top.iloc[-1], top.iloc[-2]
report.show(
    fig,
    "Paired horizontal bar charts ranking scheduling tools by how widely they are "
    "named. The left panel is the share of firms with at least one posting naming "
    "the tool; the right panel is the share of postings, with small firms' rates "
    "shrunk toward the average. "
    f"{lead['Tool']} leads on both, named by {lead['% firms naming it']:.1f} percent "
    f"of firms and {lead['% postings (shrunk)']:.1f} percent of postings, ahead of "
    f"{runner['Tool']} at {runner['% firms naming it']:.1f} and "
    f"{runner['% postings (shrunk)']:.1f} percent. The table above the figure has "
    "every value.",
)
Paired horizontal bar charts ranking scheduling tools by how widely they are named. The left panel is the share of firms with at least one posting naming the tool; the right panel is the share of postings, with small firms' rates shrunk toward the average. Apache Airflow leads on both, named by 57.9 percent of firms and 13.1 percent of postings, ahead of Dagster at 24.6 and 1.8 percent. The table above the figure has every value.

The industry splits in two#

Firm types are different populations and are not pooled. Read down the columns: the same tool can be near-universal in one half of the industry and entirely absent from the other.

report.by_firm_type(
    d["summary"],
    [
        "airflow", "dagster", "prefect",
        "control_m", "autosys", "step_functions", "slurm",
        "databricks", "snowflake",
    ],
)
firm_type bank asset_manager multistrat systematic market_maker prop_trading fintech ALL
tool
airflow 80.0 83.0 67.0 36.0 44.0 40.0 0.0 58.0
dagster 50.0 8.0 50.0 7.0 22.0 40.0 0.0 25.0
prefect 20.0 8.0 33.0 7.0 22.0 20.0 0.0 16.0
control_m 30.0 42.0 0.0 0.0 0.0 0.0 0.0 14.0
autosys 50.0 33.0 0.0 0.0 0.0 0.0 0.0 16.0
step_functions 30.0 33.0 17.0 0.0 0.0 0.0 0.0 14.0
slurm 20.0 0.0 17.0 21.0 33.0 60.0 0.0 21.0
databricks 90.0 50.0 17.0 0.0 11.0 20.0 0.0 32.0
snowflake 90.0 75.0 50.0 14.0 0.0 20.0 100.0 44.0

Three things to take from that table.

Enterprise batch schedulers are a sell-side and traditional-asset-manager phenomenon. Control-M and AutoSys appear at banks and asset managers and at no proprietary trading firm, market maker, or multi-manager fund. They are commercial, licensed, GUI-configured products from the mainframe era, run by operations teams, and they do what Airflow does: a dependency graph of jobs with schedules, retries and alerting. The dependency-graph idea is much older than Airflow, and half the industry still buys it rather than writing it.

HPC schedulers invert that split exactly. Slurm is most common at proprietary trading firms and absent from asset managers. Where banks inherited a batch-operations department, prop shops inherited a compute cluster.

Banks buy platforms. Databricks and Snowflake are near-universal at banks, common at asset managers, and thin at trading firms. A bank’s data engineering is procurement and governance; a prop shop’s is building.

Both halves are worth knowing before you choose where to work, and the tools in this course sit between them.

What the widest columns say#

Orchestration is one slice. Sorting every tool by how many firms name it puts the scheduling layer in perspective, and two of the top entries are not products at all.

top_overall = report.tools_table(d["summary"]).head(14)
top_overall[["Tool", "Layer", "Postings", "Firms", "% firms naming it"]]
Tool Layer Postings Firms % firms naming it
0 Python language 1131 53 93.0
1 SQL (any dialect) datastore 599 46 80.7
2 Data-quality / validation work (generic) data_quality 478 46 80.7
3 CI/CD (generic) ci 843 44 77.2
4 Kubernetes container 601 40 70.2
5 Docker container 477 38 66.7
6 Terraform ci 362 37 64.9
7 Java language 574 36 63.2
8 Apache Kafka streaming 322 36 63.2
9 Apache Airflow orchestrator 164 33 57.9
10 PostgreSQL datastore 162 32 56.1
11 Bash / shell language 133 32 56.1
12 C++ language 143 29 50.9
13 Apache Spark transform 297 28 49.1

Python and SQL at the top is unsurprising and reassuring: they are the prerequisites, not the subject. Below them the picture is a delivery-and- correctness stack rather than a modelling one — CI/CD language, data-quality language, containers, infrastructure-as-code — and then Airflow.

Containers deserve a note, because this course does not cover them. Kubernetes and Docker are named by more firms than any other product in the corpus, ahead of Airflow (the table above has the numbers). That is a real gap between what the data shows and what the syllabus teaches, and it is worth naming rather than hiding.

The reason for the choice: containers solve where your code runs, which is a deployment concern owned by a platform team at most of these firms, while this course is about whether your result is right and can be rebuilt. A student who can build a tested, reproducible pipeline will learn Docker in a week on the job; the reverse is not true. But if you are choosing electives, this is the strongest signal in the data for one the course does not give you.

The finding that does not depend on any tool#

Two features track whether postings discuss the work this course is about rather than naming a product: data-quality language (validation, lineage, reconciliation, testing) and reproducibility language (idempotency, point-in-time correctness, backfills, look-ahead bias).

generic = report.by_firm_type(
    d["summary"], ["data_quality_generic", "reproducibility_generic", "airflow"]
)
generic
firm_type bank asset_manager multistrat systematic market_maker prop_trading fintech ALL
tool
data_quality_generic 100.0 100.0 100.0 57.0 56.0 100.0 0.0 81.0
reproducibility_generic 80.0 42.0 50.0 21.0 44.0 40.0 100.0 46.0
airflow 80.0 83.0 67.0 36.0 44.0 40.0 0.0 58.0

Data-quality language is named by more firms than any product, and in every firm type it is at least as common as the most-named product there. (The fintech column is a single firm, so it is one board, not a population.) That is the strongest evidence in this dataset for the shape of the course: the durable skill is reasoning about whether a number is right, and the tooling is how you make that reasoning repeatable.

It is also the finding least likely to go stale. Airflow may not be the default orchestrator in five years. “Can you show this result is correct and rebuild it from raw data” will still be the question.

What this supports, and what it does not#

Worth being clear about the reach of a sample like this. It is not a randomised study of the industry and does not need to be: it is a census of what a defined set of firms chose to publish, which is a real thing to measure and is large enough that the big patterns are stable. Read with that in mind, it supports:

  • Airflow is the orchestrator most firms name, by the widest margin, on every way of counting we tried. Teaching it by name is justified.

  • The dependency-graph concept generalises across the industry’s two halves: enterprise schedulers on the sell side, HPC schedulers at prop shops, container orchestrators nearly everywhere. Teaching the idea with a small local runner and then naming the production tools matches what firms describe.

  • Correctness and automation language is more common than any single product, so leading with those and treating tools as instruments is the right ordering.

It does not support:

  • Claims about market share or about what firms actually run. This is hiring language, and a firm can run a tool for a decade without naming it.

  • Per-firm claims resting on a handful of postings.

  • Trend claims, yet. Those need several collection runs, and the dataset is built so they become available once enough exist.

The methods notebook in the project repository has the sample frame, every firm’s caveats, the four ways of counting and what each one over- and under-states, and the rules that decide what counts as a mention. If a number here looks wrong to you, that is where to check it.