Job Postings: Sample, Methods and Limitations#

The reference document behind What Firms Ask For. Everything the findings page asserts is derived here, and everything it defers is spelled out here. It reads in the order you would need it to check a number:

  1. The data: what a record is and where it comes from.

  2. The sample: which firms are in it, which are not, and why.

  3. Filters: which postings count, and what each choice costs.

  4. What counts as a mention: the rules that turn text into features.

  5. How to read the numbers: four ways of counting, and what each over- and under-states.

  6. Results: the full tables the findings page draws from.

  7. What this cannot support.

Every table below is built by the project’s doit pipeline from the committed corpus. This notebook computes nothing; it loads those tables and shows them.

from jobpostings import report

d = report.load()
m = report.load_methods(d)
report.headline(d)
collection runs: 2026-09-25, 2026-09-27
11,346 postings from 64 firms
9,237 technical postings
1,600 mention a pipeline, ETL/ELT or orchestration

The data#

One record is one job posting, keyed <firm>:<route>:<native_id>, where the route is the applicant-tracking system the firm uses to publish its own openings (Greenhouse, Workday, Lever and so on). Each record has the title, location, department and full posting text, and the run on which it was first seen.

Collection is a manual step, run every few months, and it is deliberately not part of the reproducible build. It touches dozens of firms’ careers sites, and re-running it does not reproduce the past: postings come down once roles are filled. So the collected corpus is committed to the repository (data_manual/), along with the raw API payloads it was parsed from, and everything from here on is a pure function of that committed corpus plus one rules file.

A posting seen on several runs is one record with several observations, so the same tables can later be recomputed per run. The runs so far:

m["collection_runs"]
Run date Firms Postings seen New Changed Newly closed
0 2026-09-25 28 1993 1993 0 0
1 2026-09-27 64 7655 5664 369 2

Until runs are spread over months rather than days, this is one snapshot, and no trend claim is available.

Where the language model is, and is not#

Finding each firm’s endpoint is the one fuzzy step. src/discover_sources.py asks Claude exactly one question per firm: which applicant-tracking system, and what is the endpoint? That is genuinely hard to automate: one firm’s board token is wehrtyou, another’s appears only inside JSON embedded in an unrelated page, a third’s differs from the obvious guess by two characters and the obvious guess returns an empty board.

The answer is then validated by calling the endpoint and requiring real postings back before anything is written to the manifest. Claude never reads a posting, never extracts a feature and never produces a number. Every count here comes from the declarative rules described below. The proposer does not have to be trusted, because the gate confers the trust: a hand-written proposal clears exactly the same check.

The sample#

The sample frame is data_manual/manifest.yml: firms whose careers systems expose a machine-readable endpoint we have found and validated. That is a real selection rule with real consequences.

  • It is not a random sample of quantitative finance. It over-represents firms that publish openly and under-represents firms that recruit through networks and headhunters.

  • Coverage is systematically worse for banks, which run fragmented legacy applicant-tracking systems.

  • Firms differ in board size by three orders of magnitude, from a single posting to a few thousand. That fact drives the choice of estimator below.

One row per firm that contributed postings, largest board first:

m["firms_collected"]
Firm Type Route Sampling Stored scope Postings Technical Pipeline-mentioning
0 JPMorganChase bank oracle_recruiting census technical_only 2425 2425 501
1 HSBC bank eightfold census all_postings 1472 176 38
2 Citi bank workday census technical_only 1023 1023 178
3 Goldman Sachs bank oracle_recruiting census technical_only 559 559 63
4 Barclays bank workday census technical_only 421 421 80
5 Morgan Stanley bank workday census technical_only 400 400 50
6 Deutsche Bank bank workday census technical_only 361 361 46
7 Jane Street prop_trading greenhouse census all_postings 230 150 4
8 Bank of America bank workday census technical_only 220 220 48
9 Point72 (incl. Cubist Systematic Strategies) multistrat greenhouse census all_postings 220 166 23
10 Millennium Management multistrat eightfold census all_postings 215 162 28
11 Vanguard asset_manager workday census technical_only 211 211 41
12 Qube Research & Technologies systematic greenhouse census all_postings 199 182 31
13 Fidelity Investments asset_manager workday census technical_only 192 192 36
14 Citadel Securities market_maker wayback_sitemap census all_postings 189 152 11
15 Northern Trust asset_manager workday census technical_only 181 181 26
16 IMC Trading market_maker greenhouse census all_postings 175 115 17
17 Wells Fargo bank workday census technical_only 171 171 43
18 Optiver market_maker greenhouse census all_postings 166 108 19
19 DRW prop_trading drw_next census all_postings 164 122 26
20 Citadel multistrat wayback_sitemap census all_postings 154 108 2
21 BlackRock asset_manager workday census technical_only 154 154 49
22 Jump Trading prop_trading greenhouse census all_postings 113 94 17
23 WorldQuant systematic greenhouse census all_postings 101 87 18
24 Jefferies bank oracle_recruiting census technical_only 92 92 16
25 Tower Research Capital prop_trading greenhouse census all_postings 91 74 21
26 The D. E. Shaw Group systematic deshaw census all_postings 90 42 5
27 Invesco asset_manager workday census technical_only 90 90 18
28 Squarepoint Capital systematic greenhouse census all_postings 89 58 11
29 Hudson River Trading prop_trading greenhouse census all_postings 87 62 5
30 PIMCO asset_manager workday census technical_only 84 84 3
31 Schonfeld Strategic Advisors multistrat greenhouse census all_postings 80 40 9
32 Wellington Management asset_manager workday census technical_only 76 76 4
33 Capital Group asset_manager workday census technical_only 67 67 14
34 T. Rowe Price asset_manager workday census technical_only 63 63 17
35 Man Group (incl. AHL) asset_manager greenhouse census all_postings 59 38 9
36 Two Sigma systematic avature census all_postings 53 43 9
37 The Voleon Group systematic ashby census all_postings 53 38 15
38 G-Research systematic workday census technical_only 53 53 6
39 AQR Capital Management systematic greenhouse census all_postings 51 36 4
40 Virtu Financial market_maker greenhouse census all_postings 49 36 1
41 Flow Traders market_maker greenhouse census all_postings 47 33 2
42 Akuna Capital market_maker greenhouse census all_postings 42 33 6
43 Old Mission Capital market_maker greenhouse census all_postings 37 29 1
44 Clear Street fintech greenhouse census all_postings 36 19 2
45 Dimensional Fund Advisors asset_manager workday census technical_only 30 30 6
46 Verition Fund Management multistrat greenhouse census all_postings 26 17 2
47 Wolverine Trading market_maker pinpoint census all_postings 24 19 1
48 Belvedere Trading market_maker lever census all_postings 21 13 1
49 Arrowstreet Capital systematic workday census technical_only 19 19 2
50 Five Rings prop_trading greenhouse census all_postings 16 13 0
51 Bridgewater Associates systematic greenhouse census all_postings 15 6 4
52 GSA Capital systematic greenhouse census all_postings 12 10 1
53 Brevan Howard multistrat workday census technical_only 12 12 3
54 Graham Capital Management systematic greenhouse census all_postings 9 5 1
55 PDT Partners systematic greenhouse census technical_only 9 9 3
56 Winton asset_manager greenhouse census all_postings 9 7 2
57 XTX Markets Technologies systematic greenhouse census all_postings 9 8 1
58 Headlands Technologies prop_trading greenhouse census all_postings 8 6 0
59 Radix Trading prop_trading greenhouse census technical_only 8 8 0
60 Marshall Wace systematic greenhouse census all_postings 7 6 0
61 Quadrature Capital systematic greenhouse census all_postings 4 2 0
62 ExodusPoint Capital multistrat greenhouse census all_postings 2 0 0
63 PEAK6 market_maker workday census technical_only 1 1 0

Stored scope matters when comparing the Postings column across firms. Some adapters (Workday, Oracle) filter to technical titles before storing; the rest store the whole board. So Postings is not comparable across routes, while Technical is. The analysis uses technical denominators throughout.

Firms in the frame with no data#

Firms with zero postings are kept rather than dropped, which preserves the denominator of attempted coverage. Their statuses are deliberately distinguished, because they would otherwise all read as “no data”:

  • unresolved: not found yet. Says nothing about the firm. Discovery is non-deterministic: one firm resolved on one run and came back “no machine-readable board” on the next with its endpoint working the whole time.

  • unsamplable: looked at, no route we can use. Mostly this means we have not written the adapter, not that the firm cannot be reached.

  • broken: the route works in principle and the firm’s own system is failing.

m["firms_without_data"]
Status Firm Type
0 broken Balyasny Asset Management multistrat
1 unresolved Acadian Asset Management systematic
2 unresolved BNP Paribas bank
3 unresolved Rokos Capital Management multistrat
4 unresolved State Street Global Advisors asset_manager
5 unsamplable Capula Investment Management systematic
6 unsamplable Hudson Bay Capital multistrat
7 unsamplable Susquehanna International Group market_maker
8 unsamplable UBS bank

Per-firm caveats#

Recorded in the manifest next to the firm’s endpoint and carried into the frame, so a count and its caveat cannot drift apart.

m["firm_caveats"]
Firm Status Postings Caveat
0 AQR Capital Management resolved 51 Postings name essentially no infrastructure products.
1 Arrowstreet Capital resolved 19 POST {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returned total=23 and requires offset paging (20 per page); a second site, Campus_Careers, on the same host/tenant holds 4 additional student/intern postings and must be queried separately to get full coverage.
2 Balyasny Asset Management broken 0 Workday tenant returns HTTP 401 "Unable to verify credentials for system account 39$1476" on the CXS API and serves a "Workday is currently unavailable" page for individual roles. Failure is on the firm's side, not a block on us. Retry on each run.
3 Bank of America resolved 220 POST to the CXS endpoint returns total=1950 with 20 postings per page (paginate via offset); this Workday site is experienced/lateral hiring only — campus/university roles live on a separate Avature board (bankcampuscareers.tal.net), and no other ghr site name resolved.
4 Barclays resolved 421 POST {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returns total=784 and pages correctly at offset=760; Workday caps each response at 20 postings, so full harvest needs ~40 paged requests plus per-job /job/<externalPath> calls for descriptions.
5 Belvedere Trading resolved 21 belvederetrading.com/careers links directly to jobs.lever.co/belvederetrading; the v0 postings endpoint returns all 21 live postings in one unpaginated response with full descriptions, so no sampling limitation beyond it showing only currently-published roles.
6 BlackRock resolved 154 careers.blackrock.com is a Radancy/TalentBrew front end whose experienced-hire apply links all resolve to this Workday tenant (total 315, paginates cleanly to offset 300); early-career/campus postings are the sampling gap, since those apply through blackrock.tal.net and Beamery flows instead and are not in this Workday site (no other site name on the tenant responded).
7 Brevan Howard resolved 12 POST {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returns total=28 and full pagination (offset=20 yields the remaining 8); brevanhoward.com/careers is a Phenom CareerConnect front-end (tenant BHABHAGB) but the underlying req/apply system is this Workday tenant, so a Phenom-only scrape could in principle surface postings not in this feed.
8 Bridgewater Associates resolved 15 Public Greenhouse board carries ~15 roles total, far fewer than the firm hires. Token is only discoverable in the JSON embedded at bridgewater.com/jobboard. Any per-firm rate here rests on a handful of postings.
9 Capital Group resolved 67 POST returned total=184 but only 20 per page, so full harvest requires paging offset in increments of 20 (limit caps at 20).
10 Capula Investment Management unsamplable 0 Careers pages link exclusively to Workable. Workable has a public API but the account slug could not be determined; revisit if the slug is found.
11 Citadel resolved 154 citadel.com blocks automated fetching (Cloudflare 403) for every client we tried. The live career sitemap supplies the URL list, but posting text comes from Wayback Machine snapshots, so vintages are mixed and some archived postings are no longer open. Treat Citadel counts as recent-history, not a current snapshot. Text extraction also picks up site navigation chrome, which is stripped in normalisation.
12 Citadel Securities resolved 189 [medium confidence at discovery] Citadel Securities runs no third-party ATS — the apply form is their in-house Gild system posting to WordPress admin-ajax.php with an opaque job_id UUID — and all HTML careers pages are Cloudflare-challenged (403), so the only live machine-readable source is the Yoast career-sitemap.xml (82 posting URLs, lastmod 2026-09-28), with per-posting title/location scraped through the Wayback Machine, which means individual postings may be stale or missing relative to the sitemap.
13 Citi resolved 1023 POST with {"appliedFacets":{},"limit":20,"offset":N,"searchText":""} works and reports total=2000, but offsets >=2000 silently wrap back to page 1 while the response facets sum to ~4383 postings, so full coverage requires paging per facet (e.g. Country_and_Jurisdiction); the citi.eightfold.ai/api/apply/v2/jobs?domain=citi.com alternative returns HTTP 403 "Not authorized for PCSX".
14 The D. E. Shaw Group resolved 90 Server-rendered pages, text extracted from HTML. Some listings are umbrella "all positions in X" pages rather than single roles.
15 Deutsche Bank resolved 361 POST {"appliedFacets":{},"limit":20,"offset":N,"searchText":""} returns total=1141 and paginates to offset 1120; limit is capped at 20 per request so full harvest needs ~58 paged calls, and per-job descriptions require the /wday/cxs/db/DBWebsite/job/<externalPath> detail endpoint.
16 Dimensional Fund Advisors resolved 30 POST {"appliedFacets":{},"limit":20,"offset":N,"searchText":""} returns total=70 on the first page only (later pages report total=0), so paginate by offset until a page returns fewer than limit; the official careers.dimensional.com site links to this same dimensional.wd5.myworkdayjobs.com/DFA_Careers board, and list entries give titles/locations/externalPath but need a per-job fetch for full descriptions.
17 DRW resolved 164 Listing comes from the __NEXT_DATA__ blob on the listings page; job text is scraped from each detail page's embedded props, so extraction is more fragile than a real API.
18 ExodusPoint Capital resolved 2 Very small public board; counts are not meaningful alone.
19 Fidelity Investments resolved 192 POST {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returns total=827 with 20 postings per page; tenant is fmr (FMR LLC, Fidelity's legal entity) and offset=800 still returns rows, but the API caps limit at 20 and reports total=0 on deep offsets, so pagination must loop on a non-empty jobPostings array rather than trusting total.
20 Flow Traders resolved 47 Token read directly out of a fetch() call in the HTML of https://www.flowtraders.com/careers; API returned meta.total=47 with full content for every job, so no sampling limitation.
21 Goldman Sachs resolved 559 Oracle Fusion tenant hdpc, site CX_2 (not the usual CX_1001, which returns an empty list). higher.gs.com is the marketing front end; this is the ATS behind it. Initially recorded unsamplable because our adapter had no site fallback, so the wrong site number looked like a firm with no board.
22 Graham Capital Management resolved 9 Token read directly from the API_URL constant in the JS on https://www.grahamcapital.com/careers/; endpoint returns meta.total=9 with full content, so all postings are covered in one unpaginated response.
23 G-Research resolved 53 POST with {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returns total=64 and 20 postings per page; paginate by offset (offset=60 returned the final 4, but reports total=0 on non-first pages, so trust the first page's total).
24 Headlands Technologies resolved 8 headlandstech.com/careers/ links apply buttons directly to job-boards.greenhouse.io/headlandstechnologiesllc, and the API returns all 8 postings with full content and meta.total=8, though one is a generic "Opportunistic Applications" evergreen req rather than a specific opening.
25 Hudson River Trading resolved 87 Board token is not discoverable from hudsonrivertrading.com, which only links a two-post "talent community" board (token hrttalentcommunity). The real board is wehrtyou.
26 HSBC resolved 1472 Eightfold API returns count=1467 and pages deep (start=1400 still returns global roles), but the server caps page size at 10 regardless of num, so a full pull needs ~147 sequential requests.
27 Hudson Bay Capital unsamplable 0 No applicant-tracking system at all. The careers page is a static Webflow marketing page whose only application path is a mailto: address, so there is nothing to collect.
28 Invesco resolved 90 careers.invesco.com fingerprints as myworkdayjobs; the IVZ site reports total=246 and paginates cleanly to offset 240, but a second Workday site (IVZearlycareers, tenant invesco, 6 postings) holds interns/early-career roles and must be scraped separately to get full coverage.
29 Jane Street resolved 230 Postings name languages and systems but almost no infrastructure products. The only "airflow" matches are data-centre cooling roles and are excluded by the tool rules.
30 Jefferies resolved 92 careers.jefferies.com is a Cloudflare-fronted vanity shim (returns 522 to curl) whose inline JS names the real Oracle Recruiting Cloud host hdid.fa.us2.oraclecloud.com, which serves all 167 experienced-hire postings (TotalJobsCount == items returned) in one unauthenticated GET; note that Jefferies runs campus/graduate recruiting on a separate Oleeo board at jefferies.tal.net, so student and intern programs are NOT in this feed.
31 JPMorganChase resolved 2425 jpmorganchase.com/careers links to Oracle Recruiting CE site CX_1001 on jpmc.fa.oraclecloud.com; the REST finder returns TotalJobsCount 7494 and I sampled pages at offsets 0/50/100/200/1000/5000/7450/7490 (last page returned 4 rows, confirming full pagination) rather than downloading all 7494 — note that `offset` is silently ignored unless the `facetsList` finder argument is included, which would otherwise make every page identical.
32 Marshall Wace resolved 7 mwam.com's own careers pages link to job-boards.greenhouse.io/mw-tech-grad; sampling limitation: Marshall Wace splits postings across several Greenhouse boards — mw-tech-grad (7 grad/associate roles) and mwinternshipprogram (7 internships) return postings, while the main 'marshallwace' board and 'mwam-imperial-placements' currently return 0, and experienced-hire roles appear to be agency-sourced rather than posted.
33 Millennium Management resolved 215 List endpoint omits descriptions; each posting needs a second detail call.
34 Morgan Stanley resolved 400 POST {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returns total=1345 and pages cleanly to offset 1300 (deep pages report total=0 but still return 20 postings, so paginate until an empty jobPostings array rather than trusting total); I directly sampled only 40 postings across two pages, and Morgan Stanley's morganstanley.eightfold.ai mirror returns 403 "Not authorized" so Workday is the usable endpoint.
35 Northern Trust resolved 181 careers.northerntrust.com 301-redirects to ntrs.wd1.myworkdayjobs.com/northerntrust; POST {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returns total=639 with 20 postings per page, and offset=620 still returns 19 rows, so full coverage requires paginating offset in steps of 20 (the `total` field reads 0 on non-zero-offset responses, so use the first page's total as the stop condition).
36 Optiver resolved 166 Board token is optiverus, not optiver; the obvious guess returns an empty board. Found by discovery on 2026-09-25 but missed on the 2026-09-27 sweep, so it is pinned here rather than left to chance.
37 PEAK6 resolved 1 peak6.com/careers is an Ongig embed (text-analyzer.ongig.com, group_id 1767) whose 75 listings all apply into Workday tenant peak6group split across 7 CXS sites, so the PEAK6 site alone returns only 14 (incl. the trading associate/internship reqs); scrape the siblings too for full coverage: apexfintechsolutions (49), WeInsure (5), CapMan (3), Bruce (2), ppcareers (1), TFIG (1), plus empty evilgeniuses/zogo.
38 PIMCO resolved 84 POST with {"appliedFacets":{},"limit":20,"offset":N,"searchText":""} returns total=195 on the first page and paginates cleanly to offset=180 (15 rows), but limit is capped at 20 per request and the `total` field reports 0 on non-zero offsets, so paginate until an empty page rather than trusting `total`.
39 Quadrature Capital resolved 4 Very small public board (single digits); counts are not meaningful alone.
40 Qube Research & Technologies resolved 199 qube-rt.com/careers is 403/robots-disallowed to crawlers, but real apply links resolve to job-boards.greenhouse.io/quberesearchandtechnologies and the Greenhouse API returns all 199 postings with full content in one unpaginated call (meta.total=199).
41 Radix Trading resolved 8 Radix Trading (radixtrading.co) splits postings across TWO Greenhouse boards linked from its homepage — radixuniversity (8 jobs, students/postdocs) and radixexperienced (7 jobs, https://boards-api.greenhouse.io/v1/boards/radixexperienced/jobs?content=true) — so scraping only one token samples roughly half of the 15 open roles.
42 Schonfeld Strategic Advisors resolved 80 Standardised on Prefect; their stack blurb "AWS, Prefect, Coder, Kubernetes" recurs across unrelated roles, so Prefect counts here are driven by boilerplate repetition as much as by distinct requirements.
43 Susquehanna International Group unsamplable 0 Careers site runs on iCIMS behind a Jibe front end (careers.sig.com, client code 'sig'). Endpoint exists; we have no adapter for iCIMS/Jibe.
44 T. Rowe Price resolved 63 POST returned total=138 with 20 postings per page (paginate via offset); a second site 'TRowePriceInternational' on the same host/tenant holds 29 additional non-US postings, so scraping only the TRowePrice site misses international roles.
45 Two Sigma resolved 53 Avature portal. Detail URLs ignore the slug, only the numeric id matters. Text is extracted from server-rendered HTML, not an API.
46 UBS unsamplable 0 Board runs on IBM Kenexa BrassRing Talent Gateway (jobs.ubs.com/TGnewUI, partnerid 25008, several siteids). Endpoint exists; we have no adapter for BrassRing.
47 Vanguard resolved 211 POST {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returns total=401 and 20 postings per page; www.vanguardjobs.com is a Symphony Talent/m-cloud WordPress front end whose embedded config names Workday as the ATS, and paging to offset=380 still returns rows (the `total` field resets to 0 on non-zero offsets, so page until jobPostings is empty rather than trusting it).
48 Verition Fund Management resolved 26 Token read from the Greenhouse embed script (for=veritiongroupllc) on https://www.verition.com/open-positions, whose gh_jid links match ids in the API; the single content=true call returned all 26 jobs (meta.total=26), so no pagination limitation.
49 The Voleon Group resolved 53 voleon.com/jobs embeds jobs.ashbyhq.com/voleon/embed?version=2 and the Ashby posting API returns all 53 live postings in one unpaginated response; a stale Lever board (jobs.lever.co/voleon) still exists but its API returns an empty array.
50 Wellington Management resolved 76 POST with {"appliedFacets":{},"limit":20,"offset":0,"searchText":""} returned HTTP 200 with total=129 but only 20 postings per page, so pagination via offset is required to retrieve all listings.
51 Wells Fargo resolved 171 POST with {"appliedFacets":{},"limit":20,"offset":N,"searchText":""} returns total=1705 and 20 postings per page (max limit 20); offset=1700 returns the final 5, confirming full pagination, though `total` is only populated on the offset=0 request and reads 0 on subsequent pages — the mirror host wd1.myworkdaysite.com with path /wday/cxs/wf/WellsFargoJobs/jobs returns byte-identical results.
52 Wolverine Trading resolved 24 Pinpoint board at careers.wolve.com (wolve.com is the corporate domain). Adapter added 2026-09-27 after discovery identified the platform but had no route for it.
53 XTX Markets Technologies resolved 9 Only the technology subsidiary posts to this board; XTX's main hiring is elsewhere, so the sample is a small slice of the firm.

Filters, and why#

Three decisions sit between “every posting collected” and the pool a share is computed over.

1. Technical titles only#

A posting counts as technical if its title matches a short list of stems (engineer, developer, data, quant, platform, software, scientist, analyst, research and similar), minus facilities and data-centre roles, which otherwise produce spurious “airflow” and “data” matches. The pattern is TECHNICAL_TITLE in src/jobpostings/normalize.py.

The filter excludes trading roles. That is a choice, so here is what it discards: postings whose title names a trading role and fails the filter, and the tool mentions inside them.

m["trading_exclusion"]
Measure Value
0 Postings in the corpus 11,346
1 Trading-titled postings the filter excludes 157
2 as a share of all postings 1.4%
3 Firms they come from 28
4 Tool mentions inside them 184
m["trading_exclusion_tools"]
Tool Postings Firms
0 Python 77 23
1 SQL (any dialect) 34 17
2 R 16 9
3 C++ 12 9
4 Data-quality / validation work (generic) 10 7
5 Bloomberg 7 6
6 Bash / shell 5 3
7 MATLAB 5 3
8 kdb+ / q 5 2
9 Java 3 3

What is lost is mostly languages and market data, and almost nothing from the scheduling layer this project is about.

2. Which postings form the pool#

A share is meaningless without its denominator, so every share names one of three:

  • all_technical: every technical posting. Wide; it includes low-latency C++ and trading roles that would never name an orchestrator, so it understates prevalence among the people who build pipelines.

  • pipeline_mentioning: technical postings whose text mentions a pipeline, ETL/ELT or orchestration. Much closer to the population the course is about, but selected on the outcome’s own vocabulary, so it overstates prevalence relative to engineering hiring generally. This is the default everywhere below and on the findings page.

  • pipeline_distinct_roles: the same pool with one row per (firm, title), so a role advertised in six offices counts once. The robustness check.

The last two columns show what the choice does to one tool. The share of postings naming Airflow moves several-fold between the wide and narrow pools; the share of firms naming it barely moves. That contrast is the argument for leading with firms, developed under “How to read the numbers”.

m["denominators"]
Denominator Postings Firms Airflow, % postings (pooled) Airflow, % firms naming it Definition
0 all_technical 9237 63 2.2 57.1 Every posting whose title looks technical (see TECHNICAL_TITLE in src/jobpostings/normalize.py), excluding facilities and data-centre roles. Wide, and includes many C++/trading roles that would never touch an orchestrator, so it understates tool prevalence among the people who build pipelines.
1 pipeline_mentioning 1600 57 10.2 57.9 Technical postings whose text mentions a data pipeline, ETL/ELT, or orchestration. Much closer to the population the course is about, but selected on the outcome's vocabulary, so it overstates prevalence relative to all engineering hiring.
2 pipeline_distinct_roles 1349 57 11.0 56.1 As pipeline_mentioning, but one row per (firm, title), so a role advertised in several offices counts once.

3. Distinct roles, not duplicates#

Firms post the same role separately for each office. Those are distinct requisitions, not duplicate records, so the main pools keep them all. The pipeline_distinct_roles pool above is how to check that this does not drive a result.

What counts as a mention#

Every feature comes from a declarative rule in src/tool_rules.yml. The rules carry context tests because a naive keyword search is wrong in ways that change conclusions, and they are unit-tested against the exact strings that fooled earlier versions.

The clearest case is Airflow, which is also a facilities term. At two firms in this corpus the only keyword matches were data-centre cooling roles, and counting those would have put a firm that never names an orchestrator near the top of the Airflow list. The rule as written:

report.code_block(m["rule_airflow"], "yaml")
airflow:
  display: Apache Airflow
  layer: orchestrator
  patterns:
  - \bairflow\b
  exclude_if_context:
  - cooling
  - \brack(-|\s)?level\b
  - \bracks?\b
  - CRAH|CRAC|HVAC
  - airflow management
  - liquid cooling
  - containment
  - data cent(er|re)
  - \bCDUs?\b|\bRDHx\b
  note: '"Airflow" is also a facilities term. Jane Street''s and HRT''s only matches were
    data-centre cooling roles ("redundancy topologies, containment, airflow management", "rack-level
    cooling and airflow"). Excluding those is the difference between Jane Street appearing
    to be an Airflow shop and the truth, which is that it never names an orchestrator.'

A hit counts only if none of exclude_if_context appears within the rule file’s context window (a few hundred characters either side of it). Other traps the rules guard against, each found by getting it wrong first: scalable matching Scala, array matching Ray, “temporal data” matching Temporal.io, a stack trace matching FINRA TRACE, Monte Carlo simulation matching a data-observability vendor, and Yarn the JavaScript package manager matching Hadoop YARN.

Worked evidence#

Counts are only as good as the sentences behind them. These are located with the same rule that produced the count, at most one per firm, so every excerpt is a hit that survived the context tests.

Airflow

report.quotes_for(m, "airflow")
Firm Title Excerpt
0 Bank of America Quantitative Finance Analyst …ata stores and big data environments. Hands-on expertise with data technologies and computing frameworks including but not limited to, Python, Spark, Airflow, Javascript and SQL. Ability to research new data technologies, architect novel data solutions for business problems and prove design approach throug…
1 Barclays AVP - Data Scientist & AI Engineer …be expected to work hands-on across on prem & cloud Compliance AI Platform. This requires strong proficiency in data ingestion and processing (Kafka, Airflow, Spark), cloud-based data platforms (Databricks, Snowflake, AWS S3/Azure), and SQL-based data transformation (dbt, PySpark). The candidate must demon…
2 BlackRock Data Engineer – AWS Native Data Platforms, Vice President …pelines for batch and event‑driven workloads, with a focus on reliability, scalability, and security. Develop and operate data workflows using Apache Airflow for orchestration and Python and SQL for transformation and data quality logic. Implement data transformations and models using modern analytics engi…
3 Brevan Howard Front Office Data Engineer …ations. Desirable Experience working with market data providers such as Bloomberg, Refinitiv or ICE. Experience with orchestration frameworks such as Airflow, Prefect or Dagster. Experience building internal tools or dashboards using Dash, Streamlit or similar frameworks. Experience with GoldenSource or ot…
4 Bridgewater Associates Research Associate, Equities Data Engineer …ata-aria-posinset="5" data-aria-level="1"> Familiarity with distributed processing and workflow orchestration (e.g., Spark, Airflow, or equivalents). <li data-leveltext=""…
5 Capital Group Data and AI Engineering Manager …ed Qualifications : 10+ years building full-stack data products end-to-end, hands-on across a modern lakehouse stack (Databricks, Unity Catalog, dbt, Airflow/Astronomer); Databricks Apps experience a plus Software engineering depth across multi-tier systems, with hands-on experience delivering information…

Control-M, an enterprise batch scheduler

report.quotes_for(m, "control_m")
Firm Title Excerpt
0 Citi Applications Support Sr Analyst - Asst. Vice President …nce in supporting applications built on Hadoop. Linux 4 - 6 years of experience Database Good SQL experience in any of the RDBMS. Scheduler Autosys / CONTROL-M or other schedulers will be of added advantage. Programming Languages UNIX shell scripting, Python / PERL will be of added advantage. Microsoft Stron…
1 Deutsche Bank Engineer …g subject areas: Very Good knowledge of the following technologies are needed: Java / Scala, Spark, Hadoop Hive, Workflow orchestrators like Airflow, Control-M, Composer Automation through Python/ Bash / Shell script Cloud offering (GCP preferred) Cloud services - IAAS, PAAS, SAAS Very good Knowledge about t…
2 Dimensional Fund Advisors Senior Site Reliability Engineer - Workflow Automation …rnetes), DAG authoring best practices, and multi-environment deployments Experience operating enterprise job scheduling platforms (e.g., Automic/UC4, Control-M, etc.) Strong Linux and Windows systems knowledge and comfort working in cloud environments (AWS preferred) Proficiency in Python for automation and…

Dagster, usually listed as an acceptable alternative rather than the core tool

report.quotes_for(m, "dagster")
Firm Title Excerpt
0 Barclays Data Engineer - Analyst …ood communication skills (both verbal and written) Desirable skills/Preferred Qualifications: • Experience with orchestration tools (Snowflake Tasks, Dagster, Prefect) • Experience in data visualization and data analytics infra using Tableau (or similar BI tool) will be a value add • Past experience with l…
1 BlackRock Senior Data Architect/Data Engineer, Aladdin Engineering - Vice President …s layer approach. Exposure to orchestration tools (e.g., Directed acyclic graph-based workflow orchestration framework for data and batch processing, Dagster, Prefect) and patterns for dependency management and backfills. Streaming and event-driven data experience (e.g., distributed event streaming and mes…
2 Brevan Howard Front Office Data Engineer …perience working with market data providers such as Bloomberg, Refinitiv or ICE. Experience with orchestration frameworks such as Airflow, Prefect or Dagster. Experience building internal tools or dashboards using Dash, Streamlit or similar frameworks. Experience with GoldenSource or other enterprise data…

How to read the numbers#

Because board sizes span three orders of magnitude, “what share of postings name this tool” is largely a statement about whoever posts the most. So each tool gets four summaries, and they are reported side by side.

Notation. Take one tool and one pool of postings. Firm \(i = 1, \dots, F\) has \(n_i\) postings in the pool, of which \(x_i\) name the tool, so its own rate is \(\hat p_i = x_i / n_i\).

% firms naming it (the extensive margin):

\[ \text{firms\_any} = \frac{1}{F} \sum_{i=1}^{F} \mathbf{1}\{x_i > 0\}. \]

The most assumption-light. A firm cannot inflate it by posting more, by reposting one role across offices, or by repeating a stack blurb. It cannot distinguish one mention from many. The findings page leads with it, alongside the raw counts \(\sum_i x_i\) (Postings) and \(\sum_i \mathbf{1}\{x_i > 0\}\) (Firms).

% postings (pooled), total mentions over total postings:

\[ \text{pooled} = \frac{\sum_i x_i}{\sum_i n_i} = \sum_i \frac{n_i}{\sum_j n_j}\, \hat p_i . \]

The second form shows the problem: each firm is weighted by its board size, so a firm with 2,400 postings outvotes forty firms with 50.

% postings (firm mean), equal weight per firm:

\[ \text{firm\_mean} = \frac{1}{F} \sum_i \hat p_i . \]

Immune to board size, but a two-posting firm reports 0% or 50% as though either were a measurement.

% postings (shrunk), the firm mean after pulling each firm’s rate toward the pooled rate in proportion to how little data it has. Assume each firm’s true rate is drawn from a \(\text{Beta}(\alpha, \beta)\) distribution and replace \(\hat p_i\) with its posterior mean:

\[ \text{shrunk} = \frac{1}{F} \sum_i \frac{x_i + \alpha}{n_i + \alpha + \beta} = \frac{1}{F} \sum_i \Big[\, w_i\, \hat p_i + (1 - w_i)\, \bar p \,\Big], \qquad w_i = \frac{n_i}{n_i + k}, \]

where \(\bar p\) is the pooled rate, \(k = \alpha + \beta\) is the prior’s weight in pseudo-postings, \(\alpha = \bar p\, k\) and \(\beta = (1 - \bar p)\, k\). The prior is fitted by method of moments: with \(v\) the across-firm sample variance of \(\hat p_i\),

\[ k = \frac{\bar p (1 - \bar p)}{v} - 1, \quad \text{clipped to } [2, 200], \]

falling back to \(k = 20\) when \(v\) is zero or too large for a beta distribution, which happens for tools almost nobody names. The second form of the formula is the interpretation: a firm with \(n_i \gg k\) keeps its own rate, and a firm with \(n_i \ll k\) is mostly the pooled rate. data_manual/index/firm_rates.csv holds every \(x_i\) and \(n_i\), so the shrinkage can be checked by hand.

The estimators, demonstrated#

Two firms: a small one whose every posting names the tool, and one with 1,000 postings that never does. Each row gives the small firm more postings. These are the production estimators, not a re-implementation.

m["estimator_demo"]
Small firm's postings (all name it) % firms naming it % pooled % firm mean % shrunk Prior weight k
0 1 50.0 0.10 50.0 2.43 20.0
1 10 50.0 0.99 50.0 17.01 20.0
2 100 50.0 9.09 50.0 42.51 20.0
3 1000 50.0 50.00 50.0 50.00 20.0

The firm mean always gives the small firm half the vote. Pooled almost never does. The shrunk estimate earns its way from one to the other as evidence accumulates. The extensive margin sits at 50% throughout, because it only asks whether a firm names the tool, never how loudly. (With only two firms the variance estimate is degenerate, so \(k\) is at its fallback of 20.)

One property worth stating because it is easy to assume otherwise: the shrunk estimate does not have to lie between the firm mean and the pooled rate. From the formula, it averages per-firm weights \(w_i\) that differ across firms, so when board size correlates with a firm’s own rate it can sit outside that interval. In this corpus that happens for a sizeable minority of tools, almost always by a fraction of a percentage point.

A recorded test: does the weighting hold when one firm gets huge?#

Collection was briefly capped at 250 postings per firm as a storage budget. When the cap was removed (2026-09-28), one firm’s contribution grew roughly tenfold and the corpus grew by half. That was a natural experiment on the estimators. The table records how far each estimator moved across all tools at the time; it is a historical record, not recomputed by the pipeline.

estimator

mean absolute move

max

rows moving >1pp

% postings (firm mean)

0.078pp

0.474pp

0

% postings (shrunk)

0.410pp

3.067pp

9

% postings (pooled)

0.881pp

7.319pp

15

% firms naming it

1.025pp

5.263pp

32

Equal weight per firm was essentially immune: not one row moved by a percentage point. Pooled moved as a volume-weighted statistic must, toward the new firm’s profile (Java +7.3pp, CI/CD +6.9, Kubernetes +5.1, Terraform +4.8), none of which reflects a change in the market.

The extensive margin moved most. It is immune to volume inflation of a mention already counted, but not to discovering a mention that truncation had hidden. 32 tools gained firms and none lost any: a strictly one-directional correction, so the cap had been biasing the firm count downward. Capping is now opt-in and unused.

Results#

The full tables the findings page draws from, on the pipeline_mentioning pool, sorted by how many firms name each tool.

Scheduling and orchestration, every estimator#

report.tools_table(d["summary"], layers=report.SCHEDULING_LAYERS)
Tool Layer Postings Firms Firms in pool % firms naming it % postings (shrunk) % postings (firm mean) % postings (pooled)
0 Apache Airflow orchestrator 164 33 57 57.9 13.1 13.8 10.2
1 Dagster orchestrator 23 14 57 24.6 1.8 3.7 1.4
2 Slurm hpc_scheduler 26 12 57 21.1 2.2 4.1 1.6
3 Argo Workflows / Argo CD durable_workflow 23 10 57 17.5 1.5 1.6 1.4
4 CA AutoSys enterprise_scheduler 28 9 57 15.8 1.7 0.9 1.8
5 Prefect orchestrator 13 9 57 15.8 1.1 3.5 0.8
6 AWS Step Functions durable_workflow 42 8 57 14.0 2.3 1.3 2.6
7 BMC Control-M enterprise_scheduler 25 8 57 14.0 1.8 1.6 1.6
8 Azure Data Factory durable_workflow 10 6 57 10.5 0.6 0.4 0.6
9 Temporal.io durable_workflow 20 5 57 8.8 1.1 0.4 1.2
10 Hadoop YARN hpc_scheduler 9 5 57 8.8 0.6 0.3 0.6
11 Kueue hpc_scheduler 2 2 57 3.5 0.2 0.2 0.1
12 Make / Makefiles build_runner 2 2 57 3.5 0.1 0.1 0.1
13 cron / systemd timers enterprise_scheduler 1 1 57 1.8 0.1 0.0 0.1
14 HTCondor hpc_scheduler 1 1 57 1.8 0.1 0.1 0.1
15 Luigi orchestrator 1 1 57 1.8 0.1 0.0 0.1
16 Apache NiFi durable_workflow 1 1 57 1.8 0.1 0.0 0.1
17 Tidal Workload Automation enterprise_scheduler 1 1 57 1.8 0.1 0.0 0.1

Who names no orchestrator at all#

Firms with at least 20 technical postings, none of which names Airflow, Prefect, Dagster, Luigi, Control-M or AutoSys. This is not random and not a gap in the data: latency-first proprietary shops describe building their own scheduling layer instead of adopting a product.

m["no_orchestrator_firms"]
Firm Type Technical postings
0 Jane Street prop_trading 150
1 Citadel multistrat 108
2 Jump Trading prop_trading 94
3 PIMCO asset_manager 84
4 Hudson River Trading prop_trading 62
5 Two Sigma systematic 43
6 The D. E. Shaw Group systematic 42
7 AQR Capital Management systematic 36
8 Akuna Capital market_maker 33
9 Old Mission Capital market_maker 29

Every other layer#

The full ranking within each layer, so the syllabus can be argued from the data rather than from the orchestration slice alone.

report.layer_rankings(d["summary"], report.OTHER_LAYERS)
Layer Tool Postings Firms % firms naming it % postings (shrunk)
0 transform Apache Spark 297 28 49.1 14.6
1 transform pandas 98 26 45.6 7.2
2 transform NumPy 69 20 35.1 4.8
3 transform dbt 50 14 24.6 3.3
4 transform Ray 31 11 19.3 2.6
5 transform Dask 21 11 19.3 1.6
6 transform Polars 20 11 19.3 1.6
7 transform DuckDB 5 3 5.3 0.4
8 datastore SQL (any dialect) 599 46 80.7 37.5
9 datastore PostgreSQL 162 32 56.1 12.5
10 datastore Snowflake 195 25 43.9 11.7
11 datastore MongoDB 104 19 33.3 5.6
12 datastore Redis 77 19 33.3 4.5
13 datastore Databricks 212 18 31.6 10.5
14 datastore BigQuery 32 11 19.3 2.3
15 datastore ClickHouse 13 8 14.0 1.1
16 datastore kdb+ / q 19 7 12.3 1.5
17 datastore ArcticDB 6 1 1.8 0.6
18 streaming Apache Kafka 322 36 63.2 18.7
19 streaming Apache Flink 53 14 24.6 2.9
20 streaming RabbitMQ 32 10 17.5 1.8
21 streaming AWS Kinesis 18 6 10.5 1.1
22 storage_format Apache Iceberg 66 14 24.6 3.7
23 storage_format Delta Lake 55 13 22.8 3.1
24 storage_format Apache Parquet 44 12 21.1 2.6
25 storage_format Apache Arrow 12 6 10.5 0.9
26 container Kubernetes 601 40 70.2 33.0
27 container Docker 477 38 66.7 26.6
28 container Helm 37 18 31.6 2.4
29 ci CI/CD (generic) 843 44 77.2 44.3
30 ci Terraform 362 37 64.9 18.7
31 ci Jenkins 219 25 43.9 11.1
32 ci GitHub Actions 67 17 29.8 3.7
33 ci GitLab CI 54 14 24.6 2.8
34 ci TeamCity 13 5 8.8 0.8
35 data_quality Data-quality / validation work (generic) 478 46 80.7 29.1
36 data_quality Reproducibility / versioning language (generic) 103 26 45.6 6.5
37 data_quality Great Expectations 8 5 8.8 0.5
38 mlops MLflow 31 13 22.8 1.9
39 mlops AWS SageMaker 32 10 17.5 1.9
40 mlops Feature store 19 9 15.8 1.2
41 mlops Kubeflow 14 7 12.3 0.8
42 ingestion dlt / Meltano / Singer 7 2 3.5 0.4
43 ingestion Airbyte 1 1 1.8 0.1
44 vendor_data Alternative data (generic) 28 13 22.8 2.5
45 vendor_data Bloomberg 17 13 22.8 2.9
46 vendor_data LSEG / Refinitiv 10 8 14.0 1.6
47 vendor_data ICE Data Services 8 6 10.5 1.0
48 vendor_data FactSet 4 2 3.5 0.6
49 vendor_data FINRA TRACE 1 1 1.8 0.1
50 language Python 1131 53 93.0 72.8
51 language Java 574 36 63.2 27.2
52 language Bash / shell 133 32 56.1 8.5
53 language C++ 143 29 50.9 13.5
54 language Go 139 27 47.4 8.4
55 language Rust 63 20 35.1 5.3
56 language Scala 61 20 35.1 3.6
57 language R 37 16 28.1 2.6
58 language MATLAB 1 1 1.8 0.1

What this cannot support#

  • What firms actually run. Postings measure hiring intent, not deployed infrastructure. A firm can run Airflow for a decade and never name it; a firm can name a tool it is only beginning to adopt.

  • Absence as evidence. Some firms publish their whole stack; others name almost nothing (see the caveats table). A zero is ambiguous between “does not use it” and “does not say”.

  • The industry as a whole. The frame is firms with a discoverable machine-readable board, not a random sample.

  • Per-firm claims resting on a handful of postings. The shrunk estimator exists because those rates are noise.

  • Trends. Those need several collection runs spread over time. The corpus records every observation so the same tables can be recomputed per run once enough exist.