What Firms Ask For#
The course map quotes ten job postings and matches their language to the course schedule. This page widens that from ten postings to a few thousand.
Every few months a script collects job postings from the applicant-tracking systems that quantitative finance firms use to publish their own openings, and counts which technologies those postings name. It is a small, reproducible measurement over a defined set of firms, and it is worth having for two reasons: the patterns in it do not depend on which ten postings someone happened to save, and building it is itself an example of the kind of pipeline this course is about — pull from an API, normalise, extract features, report, and keep the whole thing rebuildable.
This page is deliberately short. A companion methods notebook in the project repository carries the sample frame, the ways of counting, the extraction rules and the limitations, and every number here traces back to a table there.
from jobpostings import report
d = report.load()
report.headline(d)
collection runs: 2026-09-25, 2026-09-27
11,346 postings from 64 firms
9,237 technical postings
1,600 mention a pipeline, ETL/ELT or orchestration
Scheduling and orchestration#
The layer this course spends the most time on. Read the Firms column first: it counts how many distinct firms have at least one posting naming the tool, which is the hardest number to distort. A firm that advertises the same role in six offices, or repeats its stack description across unrelated postings, inflates the posting count and leaves the firm count alone.
Tools named by nobody are listed underneath, because an absence is a finding too.
sched = report.tools_table(d["summary"], layers=report.SCHEDULING_LAYERS)
sched[["Tool", "Layer", "Postings", "Firms", "% firms naming it", "% postings (shrunk)"]]
| Tool | Layer | Postings | Firms | % firms naming it | % postings (shrunk) | |
|---|---|---|---|---|---|---|
| 0 | Apache Airflow | orchestrator | 164 | 33 | 57.9 | 13.1 |
| 1 | Dagster | orchestrator | 23 | 14 | 24.6 | 1.8 |
| 2 | Slurm | hpc_scheduler | 26 | 12 | 21.1 | 2.2 |
| 3 | Argo Workflows / Argo CD | durable_workflow | 23 | 10 | 17.5 | 1.5 |
| 4 | CA AutoSys | enterprise_scheduler | 28 | 9 | 15.8 | 1.7 |
| 5 | Prefect | orchestrator | 13 | 9 | 15.8 | 1.1 |
| 6 | AWS Step Functions | durable_workflow | 42 | 8 | 14.0 | 2.3 |
| 7 | BMC Control-M | enterprise_scheduler | 25 | 8 | 14.0 | 1.8 |
| 8 | Azure Data Factory | durable_workflow | 10 | 6 | 10.5 | 0.6 |
| 9 | Temporal.io | durable_workflow | 20 | 5 | 8.8 | 1.1 |
| 10 | Hadoop YARN | hpc_scheduler | 9 | 5 | 8.8 | 0.6 |
| 11 | Kueue | hpc_scheduler | 2 | 2 | 3.5 | 0.2 |
| 12 | Make / Makefiles | build_runner | 2 | 2 | 3.5 | 0.1 |
| 13 | cron / systemd timers | enterprise_scheduler | 1 | 1 | 1.8 | 0.1 |
| 14 | HTCondor | hpc_scheduler | 1 | 1 | 1.8 | 0.1 |
| 15 | Luigi | orchestrator | 1 | 1 | 1.8 | 0.1 |
| 16 | Apache NiFi | durable_workflow | 1 | 1 | 1.8 | 0.1 |
| 17 | Tidal Workload Automation | enterprise_scheduler | 1 | 1 | 1.8 | 0.1 |
absent = report.tools_table(
d["summary"], layers=report.SCHEDULING_LAYERS, drop_empty=False
)
print("Scheduling tools named by no firm in the sample:")
for t in sorted(absent[absent["Postings"] == 0]["Tool"]):
print(f" {t}")
Scheduling tools named by no firm in the sample:
Flyte
GNU Parallel
IBM Spectrum LSF
Kedro
Kestra
Mage
Metaflow
Nextflow
PBS / Torque
Snakemake
pydoit
Airflow is named by more firms than every other orchestrator combined. Luigi, which appears in nearly every “top data pipeline tools” article, is named once in the entire corpus.
The two columns disagree in a way worth understanding. A tool can be named by many firms but in very few of their postings — breadth without depth. Dagster is the clearest case: second only to Airflow by firms, but named in barely more than one posting per firm, almost always as an acceptable alternative inside a list that begins with Airflow, rather than as the thing they build on.
fig, top = report.barh(
sched,
["% firms naming it", "% postings (shrunk)"],
["Share of firms naming the tool", "Share of postings (shrunk)"],
"Scheduling tools named in quant-finance job postings",
)
lead, runner = top.iloc[-1], top.iloc[-2]
report.show(
fig,
"Paired horizontal bar charts ranking scheduling tools by how widely they are "
"named. The left panel is the share of firms with at least one posting naming "
"the tool; the right panel is the share of postings, with small firms' rates "
"shrunk toward the average. "
f"{lead['Tool']} leads on both, named by {lead['% firms naming it']:.1f} percent "
f"of firms and {lead['% postings (shrunk)']:.1f} percent of postings, ahead of "
f"{runner['Tool']} at {runner['% firms naming it']:.1f} and "
f"{runner['% postings (shrunk)']:.1f} percent. The table above the figure has "
"every value.",
)
The industry splits in two#
Firm types are different populations and are not pooled. Read down the columns: the same tool can be near-universal in one half of the industry and entirely absent from the other.
report.by_firm_type(
d["summary"],
[
"airflow", "dagster", "prefect",
"control_m", "autosys", "step_functions", "slurm",
"databricks", "snowflake",
],
)
| firm_type | bank | asset_manager | multistrat | systematic | market_maker | prop_trading | fintech | ALL |
|---|---|---|---|---|---|---|---|---|
| tool | ||||||||
| airflow | 80.0 | 83.0 | 67.0 | 36.0 | 44.0 | 40.0 | 0.0 | 58.0 |
| dagster | 50.0 | 8.0 | 50.0 | 7.0 | 22.0 | 40.0 | 0.0 | 25.0 |
| prefect | 20.0 | 8.0 | 33.0 | 7.0 | 22.0 | 20.0 | 0.0 | 16.0 |
| control_m | 30.0 | 42.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 14.0 |
| autosys | 50.0 | 33.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 16.0 |
| step_functions | 30.0 | 33.0 | 17.0 | 0.0 | 0.0 | 0.0 | 0.0 | 14.0 |
| slurm | 20.0 | 0.0 | 17.0 | 21.0 | 33.0 | 60.0 | 0.0 | 21.0 |
| databricks | 90.0 | 50.0 | 17.0 | 0.0 | 11.0 | 20.0 | 0.0 | 32.0 |
| snowflake | 90.0 | 75.0 | 50.0 | 14.0 | 0.0 | 20.0 | 100.0 | 44.0 |
Three things to take from that table.
Enterprise batch schedulers are a sell-side and traditional-asset-manager phenomenon. Control-M and AutoSys appear at banks and asset managers and at no proprietary trading firm, market maker, or multi-manager fund. They are commercial, licensed, GUI-configured products from the mainframe era, run by operations teams, and they do what Airflow does: a dependency graph of jobs with schedules, retries and alerting. The dependency-graph idea is much older than Airflow, and half the industry still buys it rather than writing it.
HPC schedulers invert that split exactly. Slurm is most common at proprietary trading firms and absent from asset managers. Where banks inherited a batch-operations department, prop shops inherited a compute cluster.
Banks buy platforms. Databricks and Snowflake are near-universal at banks, common at asset managers, and thin at trading firms. A bank’s data engineering is procurement and governance; a prop shop’s is building.
Both halves are worth knowing before you choose where to work, and the tools in this course sit between them.
What the widest columns say#
Orchestration is one slice. Sorting every tool by how many firms name it puts the scheduling layer in perspective, and two of the top entries are not products at all.
top_overall = report.tools_table(d["summary"]).head(14)
top_overall[["Tool", "Layer", "Postings", "Firms", "% firms naming it"]]
| Tool | Layer | Postings | Firms | % firms naming it | |
|---|---|---|---|---|---|
| 0 | Python | language | 1131 | 53 | 93.0 |
| 1 | SQL (any dialect) | datastore | 599 | 46 | 80.7 |
| 2 | Data-quality / validation work (generic) | data_quality | 478 | 46 | 80.7 |
| 3 | CI/CD (generic) | ci | 843 | 44 | 77.2 |
| 4 | Kubernetes | container | 601 | 40 | 70.2 |
| 5 | Docker | container | 477 | 38 | 66.7 |
| 6 | Terraform | ci | 362 | 37 | 64.9 |
| 7 | Java | language | 574 | 36 | 63.2 |
| 8 | Apache Kafka | streaming | 322 | 36 | 63.2 |
| 9 | Apache Airflow | orchestrator | 164 | 33 | 57.9 |
| 10 | PostgreSQL | datastore | 162 | 32 | 56.1 |
| 11 | Bash / shell | language | 133 | 32 | 56.1 |
| 12 | C++ | language | 143 | 29 | 50.9 |
| 13 | Apache Spark | transform | 297 | 28 | 49.1 |
Python and SQL at the top is unsurprising and reassuring: they are the prerequisites, not the subject. Below them the picture is a delivery-and- correctness stack rather than a modelling one — CI/CD language, data-quality language, containers, infrastructure-as-code — and then Airflow.
Containers deserve a note, because this course does not cover them. Kubernetes and Docker are named by more firms than any other product in the corpus, ahead of Airflow (the table above has the numbers). That is a real gap between what the data shows and what the syllabus teaches, and it is worth naming rather than hiding.
The reason for the choice: containers solve where your code runs, which is a deployment concern owned by a platform team at most of these firms, while this course is about whether your result is right and can be rebuilt. A student who can build a tested, reproducible pipeline will learn Docker in a week on the job; the reverse is not true. But if you are choosing electives, this is the strongest signal in the data for one the course does not give you.
The finding that does not depend on any tool#
Two features track whether postings discuss the work this course is about rather than naming a product: data-quality language (validation, lineage, reconciliation, testing) and reproducibility language (idempotency, point-in-time correctness, backfills, look-ahead bias).
generic = report.by_firm_type(
d["summary"], ["data_quality_generic", "reproducibility_generic", "airflow"]
)
generic
| firm_type | bank | asset_manager | multistrat | systematic | market_maker | prop_trading | fintech | ALL |
|---|---|---|---|---|---|---|---|---|
| tool | ||||||||
| data_quality_generic | 100.0 | 100.0 | 100.0 | 57.0 | 56.0 | 100.0 | 0.0 | 81.0 |
| reproducibility_generic | 80.0 | 42.0 | 50.0 | 21.0 | 44.0 | 40.0 | 100.0 | 46.0 |
| airflow | 80.0 | 83.0 | 67.0 | 36.0 | 44.0 | 40.0 | 0.0 | 58.0 |
Data-quality language is named by more firms than any product, and in every firm type it is at least as common as the most-named product there. (The fintech column is a single firm, so it is one board, not a population.) That is the strongest evidence in this dataset for the shape of the course: the durable skill is reasoning about whether a number is right, and the tooling is how you make that reasoning repeatable.
It is also the finding least likely to go stale. Airflow may not be the default orchestrator in five years. “Can you show this result is correct and rebuild it from raw data” will still be the question.
What this supports, and what it does not#
Worth being clear about the reach of a sample like this. It is not a randomised study of the industry and does not need to be: it is a census of what a defined set of firms chose to publish, which is a real thing to measure and is large enough that the big patterns are stable. Read with that in mind, it supports:
Airflow is the orchestrator most firms name, by the widest margin, on every way of counting we tried. Teaching it by name is justified.
The dependency-graph concept generalises across the industry’s two halves: enterprise schedulers on the sell side, HPC schedulers at prop shops, container orchestrators nearly everywhere. Teaching the idea with a small local runner and then naming the production tools matches what firms describe.
Correctness and automation language is more common than any single product, so leading with those and treating tools as instruments is the right ordering.
It does not support:
Claims about market share or about what firms actually run. This is hiring language, and a firm can run a tool for a decade without naming it.
Per-firm claims resting on a handful of postings.
Trend claims, yet. Those need several collection runs, and the dataset is built so they become available once enough exist.
The methods notebook in the project repository has the sample frame, every firm’s caveats, the four ways of counting and what each one over- and under-states, and the rules that decide what counts as a mention. If a number here looks wrong to you, that is where to check it.