01 · Core thesis

**Agent talent in this report:** People whose primary responsibilities directly involve building, operating, researching, evaluating, productizing, or technically deploying Agent systems. The counted JDs and professionals are strongly Agent-related; the displayed group names omit the repeated Agent prefix. Agent systems use models to plan steps, call tools, and complete tasks.

**Ten job groups:** Application Engineering; Research / Model Behavior; Runtime / Platform; Deployment Engineering; Product / Design; Post-training / Robustness; Evals / Quality; Orchestration / Workflow; Solutions Architecture; and Safety / Governance.

## Agent talent is several job markets sharing a production foundation

This report is based on 259 visible Agent JDs and 513 observed current Agent professionals across 32 U.S. companies. The same ten job groups connect hiring demand with talent stock, making company priorities directly comparable.

513

Observed current Agent professionals

Publicly visible lower bound

259

Active Agent JDs

Visible hiring demand

10

Unified job groups

259 JDs · 513 people

53.2%

Talent share held by the top two companies

273 / 513

_01_ The shared Agent foundation is integration, evaluation, and reliable operation

Reliable operation, Continuous evaluation, and Real-system integration appear in 87.3%, 73.4%, and 69.5% of JDs; LangChain / LangGraph appears in only 9.3%.

_02_ Application Engineering is largest, but each company's primary gap is different

Application Engineering has 77 JDs across 18 companies. Yet 11 of Google's 22 JDs are Deployment Engineering, 8 of Salesforce's 27 are Orchestration / Workflow, and OpenAI also posts 6 Post-training / Robustness and 5 Safety / Governance roles.

_03_ Talent reservoirs and the most aggressive hirers are different company sets

Microsoft and Salesforce hold 273 / 513 observed current professionals. Scale AI and OpenAI post 68 JDs combined but have only 18 observed current professionals, revealing a clear mismatch between hiring demand and visible talent stock.

_04_ The main talent sources extend beyond AI labs into cloud, consulting, and enterprise software

Amazon / AWS is the largest visible source with 21 people; Bain feeds 6 into Decagon, while Slack and MuleSoft together feed 8 into Salesforce. Different Agent roles draw from different source ecosystems.

02 · Company bottlenecks

## The same Agent label hides very different hiring needs

The same Agent label hides different needs: OpenAI centers on safe integration into real systems, Anthropic on model-behavior evaluation, Scale AI on evaluation pipelines, Salesforce on workflow validation, and ServiceNow on state-layer operation. They share an Agent label but compete for different capability stacks.

OpenAI **33 JD**

### Safe system integration

Real-system integration covers 33 of 33 JDs, and Safety / permissions covers 32 of 33. OpenAI's outlier is that almost every Agent JD touches real-system boundaries, with demand moving beyond demo-building.

Real-system integration

100.0%33 / 33

Reliable operation

100.0%33 / 33

Safety / permissions

97.0%32 / 33

Continuous evaluation

57.6%19 / 33

Agent orchestration

33.3%11 / 33

Data source: Metix AI

Anthropic **16 JD**

### Model-behavior evaluation

Continuous evaluation covers 12 of 16 JDs, above Real-system integration at 6 of 16. This is the clearest split from OpenAI: Anthropic looks more like an evaluation, environment, and model-behavior formation.

Real-system integration

37.5%6 / 16

Reliable operation

87.5%14 / 16

Safety / permissions

68.8%11 / 16

Continuous evaluation

75.0%12 / 16

Memory / state

12.5%2 / 16

Data source: Metix AI

Scale AI **35 JD**

### Evaluation factory

Continuous evaluation covers 35 of 35 JDs, the only saturated capability among the six; Agent orchestration also covers 19 of 35, pointing to scalable evaluation and feedback pipelines.

Real-system integration

80.0%28 / 35

Reliable operation

85.7%30 / 35

Safety / permissions

57.1%20 / 35

Continuous evaluation

100.0%35 / 35

Agent orchestration

54.3%19 / 35

Data source: Metix AI

Google **22 JD**

### Cloud/product orchestration

Google's Agent orchestration covers 12 of 22 JDs, above the 42.1% market baseline; Real-system integration is only 12 of 22, unlike OpenAI's saturation.

Real-system integration

54.5%12 / 22

Reliable operation

86.4%19 / 22

Safety / permissions

63.6%14 / 22

Continuous evaluation

68.2%15 / 22

Agent orchestration

54.5%12 / 22

Data source: Metix AI

Salesforce **27 JD**

### Workflow validation

Continuous evaluation covers 26 of 27 JDs, and Real-system integration covers 24 of 27. Salesforce concentrates demand on connecting enterprise workflows and measuring them continuously.

Real-system integration

88.9%24 / 27

Reliable operation

92.6%25 / 27

Safety / permissions

48.1%13 / 27

Continuous evaluation

96.3%26 / 27

Agent orchestration

48.1%13 / 27

Memory / state

3.7%1 / 27

Data source: Metix AI

ServiceNow **27 JD**

### State-layer operation

Memory / state covers 6 of 27 JDs, 3.4 times the 6.6% market baseline. ServiceNow's difference looks closer to long-running workflows, ticket state, and platform operation.

Real-system integration

44.4%12 / 27

Reliable operation

88.9%24 / 27

Safety / permissions

40.7%11 / 27

Continuous evaluation

44.4%12 / 27

Agent orchestration

33.3%9 / 27

Memory / state

22.2%6 / 27

Data source: Metix AI

Vertical companies reveal narrower capability gaps

LangChain

### State-layer company

Only 17 of 259 market JDs mention Memory / state; LangChain reaches 5 of 6, while Agent orchestration reaches 6 of 6.

**5 / 6** Memory / state**6 / 6** Agent orchestration

Snowflake

### Data-system integration

All 6 Snowflake JDs mention Real-system integration, Continuous evaluation, and Reliable operation, concentrating demand on connecting Agents to data systems and keeping them stable.

**6 / 6** Real-system integration**6 / 6** Continuous evaluation

Harvey

### Legal-workflow deployment

All 6 JDs saturate Real-system integration, Continuous evaluation, and Reliable operation, pointing to production deployment inside a vertical workflow.

**6 / 6** Real-system integration**6 / 6** Reliable operation

Databricks

### Platform operation first

Reliable operation covers 8 of 8 JDs, and Safety / permissions covers 6 of 8, tying Agent demand closely to the operating boundary of a data platform.

**8 / 8** Reliable operation**6 / 8** Safety / permissions

03 · Ten-group demand map

## Applications drive the most volume; company-specific bottlenecks create the real differentiation

The ten job groups form the Agent demand map. Application Engineering has 77 JDs, Research / Model Behavior 50, and Runtime / Platform 34, together accounting for 62.2%; company-level priorities, however, do not simply follow the market total.

### Visible hiring demand across ten job groups

Application Engineering is the largest demand group, but the most postings do not necessarily imply the greatest scarcity. A later section aligns the same job-group framework with 513 observed professionals.

Application Engineering

77 JD · 18 companies

Research / Model Behavior

50 JD · 11 companies

Runtime / Platform

34 JD · 11 companies

Deployment Engineering

19 JD · 5 companies

Product / Design

16 JD · 9 companies

Post-training / Robustness

14 JD · 4 companies

Evals / Quality

13 JD · 5 companies

Orchestration / Workflow

13 JD · 2 companies

Solutions Architecture

12 JD · 5 companies

Safety / Governance

11 JD · 5 companies

Data source: Metix AI

### Ten-group hiring mix across four focus companies

The complete JD base reveals four hiring recipes. OpenAI's 33 roles cover 8 of the 10 job groups, led by Application Engineering (10) while also investing in Post-training / Robustness (6) and Safety / Governance (5). Anthropic has 6 Research / Model Behavior roles among 16. Scale AI has 11 Application Engineering and 9 Research / Model Behavior roles among 35. Google assigns 11 of 22 roles to Deployment Engineering.

#### OpenAI

**33** JD

- Application Engineering **10**
- Research / Model Behavior **4**
- Runtime / Platform **3**
- Post-training / Robustness **6**
- Safety / Governance **5**

#### Anthropic

**16** JD

- Application Engineering **3**
- Research / Model Behavior **6**
- Runtime / Platform **3**
- Post-training / Robustness **1**
- Evals / Quality **3**

#### Scale AI

**35** JD

- Application Engineering **11**
- Research / Model Behavior **9**
- Product / Design **3**
- Post-training / Robustness **6**
- Evals / Quality **6**

#### Google

**22** JD

- Application Engineering **4**
- Runtime / Platform **3**
- Deployment Engineering **11**
- Product / Design **2**
- Safety / Governance **2**

Data source: Metix AI

04 · JD competition groups

## 259 JDs form ten competition lanes

The ten competition lanes correspond to ten primary deliverables. Companies compete for candidates within the same work object; roles carrying the Agent label but tied to different work objects have different competition boundaries. Capability terms explain bottleneck differences without changing the competition boundary.

01

### Application Engineering

Builds user-facing or domain-specific Agent products, features, and applications.

**77** JDs**18** companies

Scale AI **11**OpenAI **10**Salesforce **9**ServiceNow **7**Harvey **6**Databricks **4**Decagon **4**Google **4**Microsoft **4**Anthropic **3**Perplexity **3**Snowflake **3**Abridge **2**Hippocratic AI **2**Writer **2**Adobe **1**Cognition **1**LangChain **1**

View representative roles

- Senior / Staff Software Engineer, Developer Experience
- Software Engineer - Early Career
- Senior Engineering Manager – Agentic Product
- Senior Staff Software Engineer, API

02

### Research / Model Behavior

Researches Agent capabilities, model behavior, retrieval, planning, and applied model systems.

**50** JDs**11** companies

ServiceNow **10**Scale AI **9**Meta **7**Anthropic **6**Microsoft **6**OpenAI **4**Adobe **2**Salesforce **2**Snowflake **2**Databricks **1**Hebbia **1**

View representative roles

- Senior Machine Learning Engineer
- Staff Agentic ML Engineer - Photoshop
- Model Performance Software Engineer, Claude Code
- Research Engineer, Computer Use

03

### Runtime / Platform

Builds runtimes, harnesses, state, tool use, observability, and platform infrastructure.

**34** JDs**11** companies

ServiceNow **9**Adobe **6**Anthropic **3**Google **3**OpenAI **3**CrewAI **2**Glean **2**Hebbia **2**Salesforce **2**Databricks **1**Replit **1**

View representative roles

- Senior Software Engineer, Meta Factory Agent Harness
- Software Engineer — Model & Partnerships Management (Agentic Builders Experience team)
- Chief Architect for Adobe Intelligence Platform
- Senior AI Platform Engineer- Data and Systems

04

### Deployment Engineering

Integrates, launches, operates, and troubleshoots Agents in customer environments.

**19** JDs**5** companies

Google **11**Decagon **3**Hippocratic AI **2**OpenAI **2**Replit **1**

View representative roles

- Customer Engineer, Agent Builder
- Director of Customer Engineering, Agent Builder
- Forward Deployed Engineer III, Generative AI, Google Cloud
- Forward Deployed Engineer IV, Applied AI, Google Cloud

05

### Product / Design

Owns Agent product definition, interaction experience, design, and product direction.

**16** JDs**9** companies

Microsoft **3**Scale AI **3**Decagon **2**Google **2**Salesforce **2**Adobe **1**CrewAI **1**EvenUp **1**OpenAI **1**

View representative roles

- Principal Product Manager
- Product Manager, Agent Experience
- Senior Agent Product Manager
- Senior/Staff/Principal Product Manager - AI Orchestration & Agentic Workflows

06

### Post-training / Robustness

Improves Agent behavior through post-training, reinforcement learning, rewards, and robustness.

**14** JDs**4** companies

OpenAI **6**Scale AI **6**Anthropic **1**Snowflake **1**

View representative roles

- Research Engineer, Machine Learning (Reinforcement Learning)
- Agent Post-Training, Frontier Evals and Environments Research
- Agent Post-Training Research
- Agent Post-Training, API & Power Users

07

### Evals / Quality

Builds Agent evaluations, benchmarks, quality measurement, and feedback loops.

**13** JDs**5** companies

Scale AI **6**Anthropic **3**OpenAI **2**Databricks **1**ServiceNow **1**

View representative roles

- Data Operations Manager, Human Data
- Engineering Manager, Agent Prompts & Evals
- Staff Software Engineer - Agent Quality
- Backend Software Engineer (Evals)

08

### Orchestration / Workflow

Builds multi-step workflows, planning, routing, state, and multi-Agent coordination.

**13** JDs**2** companies

Salesforce **8**Decagon **5**

View representative roles

- Engineering Manager, Agent Orchestration
- Senior Software Engineer, Agent Orchestration
- Staff Software Engineer, Agent Orchestration
- Senior/Lead AI Software Engineer, Agentforce for Supply Chain

09

### Solutions Architecture

Designs enterprise Agent solutions, integration architecture, cloud platforms, and delivery paths.

**12** JDs**5** companies

LangChain **5**Meta **3**Salesforce **2**Databricks **1**Microsoft **1**

View representative roles

- Specialist Solutions Architect - AI/ML
- Solutions Architect (Austin)
- Solutions Architect (Dallas)
- Solutions Architect (NYC)

10

### Safety / Governance

Owns permissions, guardrails, secure execution, risk controls, and governance.

**11** JDs**5** companies

OpenAI **5**Google **2**Salesforce **2**Glean **1**Microsoft **1**

View representative roles

- Product Manager, Agent Security & Governance
- Agentic Safety and Ecosystem Architect, Trust and Safety
- Principal Product Manager, Agent 365 Security & Governance
- Principal Software Engineer, Codex Cyber

05 · Shared foundation

## Frameworks are no longer the main story: Agent hiring pays for systems that connect, evaluate, and run reliably

LangChain / LangGraph is the most-mentioned named framework, yet it appears in only 24 / 259 JDs. Reliable operation, Continuous evaluation, and Real-system integration appear in 226, 190, and 180 JDs—a 7.5–9.4× coverage gap.

### Production capabilities are far more common than framework keywords

Frameworks remain implementation paths, but no longer represent the shared requirement for Agent talent. Operation, evaluation, and integration are the capabilities that consistently recur across companies.

#### Cross-company production capabilities

These are multi-label capability signals measured across all ten job groups

Reliable operation

87.3% · 226

Continuous evaluation

73.4% · 190

Real-system integration

69.5% · 180

#### Named frameworks

The highest reaches only 24 JDs

LangChain / LangGraph

9.3% · 24

CrewAI

2.7% · 7

AutoGen

1.9% · 5

LlamaIndex

1.9% · 5

Data source: Metix AI

06 · Stock and scarcity

## Talent reservoirs and the most aggressive hirers are different company sets

Microsoft and Salesforce together hold 273 / 513 observed Agent professionals, or 53.2%. Scale AI and OpenAI have 35 and 33 active JDs but only 8 and 10 observed current professionals. Within the same job-group framework, Research / Model Behavior is the clearest scarcity signal.

### Observed Agent talent by company

Microsoft and Salesforce are the main reservoirs; high-demand OpenAI and Scale AI depend more on the external candidate market.

Microsoft

139 · 27.1%

Salesforce

134 · 26.1%

Google

44 · 8.6%

Meta

36 · 7.0%

Decagon

33 · 6.4%

ServiceNow

31 · 6.0%

Adobe

15 · 2.9%

Hippocratic AI

14 · 2.7%

Sierra

11 · 2.1%

OpenAI

10 · 1.9%

Data source: Metix AI

### Active postings per 100 observed professionals

Research / Model Behavior reaches 138.9 (50 postings / 36 professionals), the clearest scarcity signal. Safety / Governance reaches 157.1 (11 / 7) and Post-training / Robustness 140.0 (14 / 10), also showing earlier signs of pressure.

Safety / Governance

157.1 · 11 / 7

Post-training / Robustness

140.0 · 14 / 10

Research / Model Behavior

138.9 · 50 / 36

Evals / Quality

100.0 · 13 / 13

Application Engineering

98.7 · 77 / 78

Runtime / Platform

94.4 · 34 / 36

Solutions Architecture

38.7 · 12 / 31

Deployment Engineering

27.5 · 19 / 69

Orchestration / Workflow

24.5 · 13 / 53

Product / Design

8.9 · 16 / 180

Data source: Metix AI

City concentration

San Francisco

It holds 133 observed professionals (25.9%), while demand is more concentrated: San Francisco has 91 postings (35.1%) and the Bay Area totals 58.3%. New York also shows 15.1% of postings versus 5.3% of talent.

Job-group stock

35.1%

Product / Design is the largest observed stock (180 people), followed by Application Engineering (78) and Deployment Engineering (69). Public profile visibility differs by job group, so demand-stock ratios are more useful for prioritization.

07 · Sources and flows

## The sourcing map starts beyond AI labs

An external prior employer is observable for 408 professionals. Amazon / AWS is the largest source with 21 people; Bain feeds 6 into Decagon, while Slack and MuleSoft together feed 8 into Salesforce. Cloud platforms, consulting delivery, and enterprise software ecosystems supply different capabilities.

### Leading prior employers before the current Agent team

Amazon / AWS contributes 2.3 times the visible source volume of second-ranked Microsoft. Meta / Instagram, Google, LinkedIn, and Bain show why sourcing should extend beyond AI labs.

Amazon / AWS

21

Microsoft

9

Meta / Instagram

8

LinkedIn

6

Bain & Company

6

Google

6

Slack

4

Cisco

4

MuleSoft

4

Data source: Metix AI

### Six clearest talent routes

**Amazon / AWS → Microsoft** Cloud platform → enterprise AI platform

**6**

**Bain & Company → Decagon** Consulting delivery → customer-workflow Agent

**6**

**Amazon / AWS → Google** Cloud platform → orchestration and deployment

**5**

**Slack + MuleSoft → Salesforce** Enterprise software ecosystem → Agent platform

**8**

**Meta / Instagram → Google** Movement between large technology platforms

**4**

**Cisco + Deloitte + LinkedIn → Salesforce** Integration, delivery, and platform experience converges

**9**

Only the latest observable external-employer move is shown; professionals without sufficiently complete histories are excluded from routes.

01

### Fix the JD competition group first

First identify the primary deliverable within the unified ten-group taxonomy, then build the company list and compare companies and candidate pools inside that group.

02

### Then expand evidence around the bottleneck

Expand the search with evidence of operation, evaluation, integration, permissions, state, and domain delivery; a single framework name narrows the pool too early.

03

### Choose the source ecosystem last

Prioritize cloud platforms for runtime, consulting and delivery for customer workflows, and enterprise software ecosystems for enterprise Agent platforms.

08 · Notes

## This report focuses on publicly visible Agent job and talent evidence

### Coverage notes

- The report uses 259 publicly visible Agent postings to observe company hiring priorities.
- The report uses 513 observed current Agent professionals to study talent stock, sources, and movement.
- Job groups: the same ten job groups connect hiring demand with talent stock, making company priorities comparable.
- Flows: the latest prior external employer is observable for 408 people, covering 79.5% of the observed talent pool.

### Job-group definitions

- The ten job groups cover Application Engineering, Research / Model Behavior, Runtime / Platform, Deployment Engineering, Product / Design, Post-training / Robustness, Evals / Quality, Orchestration / Workflow, Solutions Architecture, and Safety / Governance.
- Job groups prioritize primary work object and responsibilities, with company brand, title keywords, and location as supporting evidence.
- Capabilities, frameworks, and domain terms explain company bottlenecks; candidate competition remains bounded by the primary job group.

### Interpretation boundaries

- 513 is a publicly visible lower bound, not the full Agent headcount at the covered companies; public profile completeness differs by job group.
- Posting counts represent visible hiring demand, not budgets, hires, or net expansion.
- Companies or job groups with fewer visible jobs or professionals are used to describe hiring focus, not to independently establish talent-competition intensity.
- Demand-stock pressure compares job-group priorities within this report coverage and should not be read as a market-wide vacancy rate.

### Questions answered

- How much current Agent talent is observed and where it concentrates by company, city, and job group.
- Total demand across the ten job groups and each company's primary hiring focus.
- Where hiring demand and current talent stock are mismatched by job group.
- Which employers are the main talent sources and what movement routes are visible.
