AI Agent talent structure | Metix AI

01 · Core thesis

Agent talent in this report: People whose primary responsibilities directly involve building, operating, researching, evaluating, productizing, or technically deploying Agent systems. The counted JDs and professionals are strongly Agent-related; the displayed group names omit the repeated Agent prefix. Agent systems use models to plan steps, call tools, and complete tasks.

Ten job groups: Application Engineering; Research / Model Behavior; Runtime / Platform; Deployment Engineering; Product / Design; Post-training / Robustness; Evals / Quality; Orchestration / Workflow; Solutions Architecture; and Safety / Governance.

Agent talent is several job markets sharing a production foundation

This report is based on 259 visible Agent JDs and 513 observed current Agent professionals across 32 U.S. companies. The same ten job groups connect hiring demand with talent stock, making company priorities directly comparable.

513

Observed current Agent professionals

Publicly visible lower bound

259

Active Agent JDs

Visible hiring demand

10

Unified job groups

259 JDs · 513 people

53.2%

Talent share held by the top two companies

273 / 513

01 The shared Agent foundation is integration, evaluation, and reliable operation

Reliable operation, Continuous evaluation, and Real-system integration appear in 87.3%, 73.4%, and 69.5% of JDs; LangChain / LangGraph appears in only 9.3%.

02 Application Engineering is largest, but each company's primary gap is different

Application Engineering has 77 JDs across 18 companies. Yet 11 of Google's 22 JDs are Deployment Engineering, 8 of Salesforce's 27 are Orchestration / Workflow, and OpenAI also posts 6 Post-training / Robustness and 5 Safety / Governance roles.

03 Talent reservoirs and the most aggressive hirers are different company sets

Microsoft and Salesforce hold 273 / 513 observed current professionals. Scale AI and OpenAI post 68 JDs combined but have only 18 observed current professionals, revealing a clear mismatch between hiring demand and visible talent stock.

04 The main talent sources extend beyond AI labs into cloud, consulting, and enterprise software

Amazon / AWS is the largest visible source with 21 people; Bain feeds 6 into Decagon, while Slack and MuleSoft together feed 8 into Salesforce. Different Agent roles draw from different source ecosystems.

02 · Company bottlenecks

The same Agent label hides very different hiring needs

The same Agent label hides different needs: OpenAI centers on safe integration into real systems, Anthropic on model-behavior evaluation, Scale AI on evaluation pipelines, Salesforce on workflow validation, and ServiceNow on state-layer operation. They share an Agent label but compete for different capability stacks.

OpenAI 33 JD

Safe system integration

Real-system integration covers 33 of 33 JDs, and Safety / permissions covers 32 of 33. OpenAI's outlier is that almost every Agent JD touches real-system boundaries, with demand moving beyond demo-building.

Real-system integration

100.0%33 / 33

Reliable operation

100.0%33 / 33

Safety / permissions

97.0%32 / 33

Continuous evaluation

57.6%19 / 33

Agent orchestration

33.3%11 / 33

Data source: Metix AI

Anthropic 16 JD

Model-behavior evaluation

Continuous evaluation covers 12 of 16 JDs, above Real-system integration at 6 of 16. This is the clearest split from OpenAI: Anthropic looks more like an evaluation, environment, and model-behavior formation.

Real-system integration

37.5%6 / 16

Reliable operation

87.5%14 / 16

Safety / permissions

68.8%11 / 16

Continuous evaluation

75.0%12 / 16

Memory / state

12.5%2 / 16

Data source: Metix AI

Scale AI 35 JD

Evaluation factory

Continuous evaluation covers 35 of 35 JDs, the only saturated capability among the six; Agent orchestration also covers 19 of 35, pointing to scalable evaluation and feedback pipelines.

Real-system integration

80.0%28 / 35

Reliable operation

85.7%30 / 35

Safety / permissions

57.1%20 / 35

Continuous evaluation

100.0%35 / 35

Agent orchestration

54.3%19 / 35

Data source: Metix AI

Google 22 JD

Cloud/product orchestration

Google's Agent orchestration covers 12 of 22 JDs, above the 42.1% market baseline; Real-system integration is only 12 of 22, unlike OpenAI's saturation.

Real-system integration

54.5%12 / 22

Reliable operation

86.4%19 / 22

Safety / permissions

63.6%14 / 22

Continuous evaluation

68.2%15 / 22

Agent orchestration

54.5%12 / 22

Data source: Metix AI

Salesforce 27 JD

Workflow validation

Continuous evaluation covers 26 of 27 JDs, and Real-system integration covers 24 of 27. Salesforce concentrates demand on connecting enterprise workflows and measuring them continuously.

Real-system integration

88.9%24 / 27

Reliable operation

92.6%25 / 27

Safety / permissions

48.1%13 / 27

Continuous evaluation

96.3%26 / 27

Agent orchestration

48.1%13 / 27

Memory / state

3.7%1 / 27

Data source: Metix AI

ServiceNow 27 JD

State-layer operation

Memory / state covers 6 of 27 JDs, 3.4 times the 6.6% market baseline. ServiceNow's difference looks closer to long-running workflows, ticket state, and platform operation.

Real-system integration

44.4%12 / 27

Reliable operation

88.9%24 / 27

Safety / permissions

40.7%11 / 27

Continuous evaluation

44.4%12 / 27

Agent orchestration

33.3%9 / 27

Memory / state

22.2%6 / 27

Data source: Metix AI

Vertical companies reveal narrower capability gaps

LangChain

State-layer company

Only 17 of 259 market JDs mention Memory / state; LangChain reaches 5 of 6, while Agent orchestration reaches 6 of 6.

5 / 6 Memory / state6 / 6 Agent orchestration

Snowflake

Data-system integration

All 6 Snowflake JDs mention Real-system integration, Continuous evaluation, and Reliable operation, concentrating demand on connecting Agents to data systems and keeping them stable.

6 / 6 Real-system integration6 / 6 Continuous evaluation

Harvey

Legal-workflow deployment

All 6 JDs saturate Real-system integration, Continuous evaluation, and Reliable operation, pointing to production deployment inside a vertical workflow.

6 / 6 Real-system integration6 / 6 Reliable operation

Databricks

Platform operation first

Reliable operation covers 8 of 8 JDs, and Safety / permissions covers 6 of 8, tying Agent demand closely to the operating boundary of a data platform.

8 / 8 Reliable operation6 / 8 Safety / permissions

03 · Ten-group demand map

Applications drive the most volume; company-specific bottlenecks create the real differentiation

The ten job groups form the Agent demand map. Application Engineering has 77 JDs, Research / Model Behavior 50, and Runtime / Platform 34, together accounting for 62.2%; company-level priorities, however, do not simply follow the market total.

Visible hiring demand across ten job groups

Application Engineering is the largest demand group, but the most postings do not necessarily imply the greatest scarcity. A later section aligns the same job-group framework with 513 observed professionals.

Application Engineering

77 JD · 18 companies

Research / Model Behavior

50 JD · 11 companies

Runtime / Platform

34 JD · 11 companies

Deployment Engineering

19 JD · 5 companies

Product / Design

16 JD · 9 companies

Post-training / Robustness

14 JD · 4 companies

Evals / Quality

13 JD · 5 companies

Orchestration / Workflow

13 JD · 2 companies

Solutions Architecture

12 JD · 5 companies

Safety / Governance

11 JD · 5 companies

Data source: Metix AI

Ten-group hiring mix across four focus companies

The complete JD base reveals four hiring recipes. OpenAI's 33 roles cover 8 of the 10 job groups, led by Application Engineering (10) while also investing in Post-training / Robustness (6) and Safety / Governance (5). Anthropic has 6 Research / Model Behavior roles among 16. Scale AI has 11 Application Engineering and 9 Research / Model Behavior roles among 35. Google assigns 11 of 22 roles to Deployment Engineering.

OpenAI

33 JD

Anthropic

16 JD

Scale AI

35 JD

Google

22 JD

Data source: Metix AI

04 · JD competition groups

259 JDs form ten competition lanes

The ten competition lanes correspond to ten primary deliverables. Companies compete for candidates within the same work object; roles carrying the Agent label but tied to different work objects have different competition boundaries. Capability terms explain bottleneck differences without changing the competition boundary.

01

Application Engineering

Builds user-facing or domain-specific Agent products, features, and applications.

77 JDs18 companies

Scale AI 11OpenAI 10Salesforce 9ServiceNow 7Harvey 6Databricks 4Decagon 4Google 4Microsoft 4Anthropic 3Perplexity 3Snowflake 3Abridge 2Hippocratic AI 2Writer 2Adobe 1Cognition 1LangChain 1

View representative roles

02

Research / Model Behavior

Researches Agent capabilities, model behavior, retrieval, planning, and applied model systems.

50 JDs11 companies

ServiceNow 10Scale AI 9Meta 7Anthropic 6Microsoft 6OpenAI 4Adobe 2Salesforce 2Snowflake 2Databricks 1Hebbia 1

View representative roles

03

Runtime / Platform

Builds runtimes, harnesses, state, tool use, observability, and platform infrastructure.

34 JDs11 companies

ServiceNow 9Adobe 6Anthropic 3Google 3OpenAI 3CrewAI 2Glean 2Hebbia 2Salesforce 2Databricks 1Replit 1

View representative roles

04

Deployment Engineering

Integrates, launches, operates, and troubleshoots Agents in customer environments.

19 JDs5 companies

Google 11Decagon 3Hippocratic AI 2OpenAI 2Replit 1

View representative roles

05

Product / Design

Owns Agent product definition, interaction experience, design, and product direction.

16 JDs9 companies

Microsoft 3Scale AI 3Decagon 2Google 2Salesforce 2Adobe 1CrewAI 1EvenUp 1OpenAI 1

View representative roles

06

Post-training / Robustness

Improves Agent behavior through post-training, reinforcement learning, rewards, and robustness.

14 JDs4 companies

OpenAI 6Scale AI 6Anthropic 1Snowflake 1

View representative roles

07

Evals / Quality

Builds Agent evaluations, benchmarks, quality measurement, and feedback loops.

13 JDs5 companies

Scale AI 6Anthropic 3OpenAI 2Databricks 1ServiceNow 1

View representative roles

08

Orchestration / Workflow

Builds multi-step workflows, planning, routing, state, and multi-Agent coordination.

13 JDs2 companies

Salesforce 8Decagon 5

View representative roles

09

Solutions Architecture

Designs enterprise Agent solutions, integration architecture, cloud platforms, and delivery paths.

12 JDs5 companies

LangChain 5Meta 3Salesforce 2Databricks 1Microsoft 1

View representative roles

10

Safety / Governance

Owns permissions, guardrails, secure execution, risk controls, and governance.

11 JDs5 companies

OpenAI 5Google 2Salesforce 2Glean 1Microsoft 1

View representative roles

05 · Shared foundation

Frameworks are no longer the main story: Agent hiring pays for systems that connect, evaluate, and run reliably

LangChain / LangGraph is the most-mentioned named framework, yet it appears in only 24 / 259 JDs. Reliable operation, Continuous evaluation, and Real-system integration appear in 226, 190, and 180 JDs—a 7.5–9.4× coverage gap.

Production capabilities are far more common than framework keywords

Frameworks remain implementation paths, but no longer represent the shared requirement for Agent talent. Operation, evaluation, and integration are the capabilities that consistently recur across companies.

Cross-company production capabilities

These are multi-label capability signals measured across all ten job groups

Reliable operation

87.3% · 226

Continuous evaluation

73.4% · 190

Real-system integration

69.5% · 180

Named frameworks

The highest reaches only 24 JDs

LangChain / LangGraph

9.3% · 24

CrewAI

2.7% · 7

AutoGen

1.9% · 5

LlamaIndex

1.9% · 5

Data source: Metix AI

06 · Stock and scarcity

Talent reservoirs and the most aggressive hirers are different company sets

Microsoft and Salesforce together hold 273 / 513 observed Agent professionals, or 53.2%. Scale AI and OpenAI have 35 and 33 active JDs but only 8 and 10 observed current professionals. Within the same job-group framework, Research / Model Behavior is the clearest scarcity signal.

Observed Agent talent by company

Microsoft and Salesforce are the main reservoirs; high-demand OpenAI and Scale AI depend more on the external candidate market.

Microsoft

139 · 27.1%

Salesforce

134 · 26.1%

Google

44 · 8.6%

Meta

36 · 7.0%

Decagon

33 · 6.4%

ServiceNow

31 · 6.0%

Adobe

15 · 2.9%

Hippocratic AI

14 · 2.7%

Sierra

11 · 2.1%

OpenAI

10 · 1.9%

Data source: Metix AI

Active postings per 100 observed professionals

Research / Model Behavior reaches 138.9 (50 postings / 36 professionals), the clearest scarcity signal. Safety / Governance reaches 157.1 (11 / 7) and Post-training / Robustness 140.0 (14 / 10), also showing earlier signs of pressure.

Safety / Governance

157.1 · 11 / 7

Post-training / Robustness

140.0 · 14 / 10

Research / Model Behavior

138.9 · 50 / 36

Evals / Quality

100.0 · 13 / 13

Application Engineering

98.7 · 77 / 78

Runtime / Platform

94.4 · 34 / 36

Solutions Architecture

38.7 · 12 / 31

Deployment Engineering

27.5 · 19 / 69

Orchestration / Workflow

24.5 · 13 / 53

Product / Design

8.9 · 16 / 180

Data source: Metix AI

City concentration

San Francisco

It holds 133 observed professionals (25.9%), while demand is more concentrated: San Francisco has 91 postings (35.1%) and the Bay Area totals 58.3%. New York also shows 15.1% of postings versus 5.3% of talent.

Job-group stock

35.1%

Product / Design is the largest observed stock (180 people), followed by Application Engineering (78) and Deployment Engineering (69). Public profile visibility differs by job group, so demand-stock ratios are more useful for prioritization.

07 · Sources and flows

The sourcing map starts beyond AI labs

An external prior employer is observable for 408 professionals. Amazon / AWS is the largest source with 21 people; Bain feeds 6 into Decagon, while Slack and MuleSoft together feed 8 into Salesforce. Cloud platforms, consulting delivery, and enterprise software ecosystems supply different capabilities.

Leading prior employers before the current Agent team

Amazon / AWS contributes 2.3 times the visible source volume of second-ranked Microsoft. Meta / Instagram, Google, LinkedIn, and Bain show why sourcing should extend beyond AI labs.

Amazon / AWS

21

Microsoft

9

Meta / Instagram

8

LinkedIn

6

Bain & Company

6

Google

6

Slack

4

Cisco

4

MuleSoft

4

Data source: Metix AI

Six clearest talent routes

Amazon / AWS → Microsoft Cloud platform → enterprise AI platform

6

Bain & Company → Decagon Consulting delivery → customer-workflow Agent

6

Amazon / AWS → Google Cloud platform → orchestration and deployment

5

Slack + MuleSoft → Salesforce Enterprise software ecosystem → Agent platform

8

Meta / Instagram → Google Movement between large technology platforms

4

Cisco + Deloitte + LinkedIn → Salesforce Integration, delivery, and platform experience converges

9

Only the latest observable external-employer move is shown; professionals without sufficiently complete histories are excluded from routes.

01

Fix the JD competition group first

First identify the primary deliverable within the unified ten-group taxonomy, then build the company list and compare companies and candidate pools inside that group.

02

Then expand evidence around the bottleneck

Expand the search with evidence of operation, evaluation, integration, permissions, state, and domain delivery; a single framework name narrows the pool too early.

03

Choose the source ecosystem last

Prioritize cloud platforms for runtime, consulting and delivery for customer workflows, and enterprise software ecosystems for enterprise Agent platforms.

08 · Notes

This report focuses on publicly visible Agent job and talent evidence

Coverage notes

Job-group definitions

Interpretation boundaries

Questions answered