ai-infra-jobs

AI infrastructure roles, filterable by the stack you actually work on.

Distributed training · inference serving · GPU fleets · network fabric — aggregated straight from company boards, never a copy of a copy.

883 open roles · 17 companies · last verified today

Baseteninference provider

Montreal · New York +2 more · remote · $165K–$330K · senior

distributed-inferencenetwork-fabricnvidiacollectivesgpu-kernels

posted 5mo ago · verified today

Baseteninference provider

San Francisco · remote · $200K–$275K · senior

post-trainingresearch-engineertraining-frameworksfine-tuninggpu-generic

posted 4mo ago · verified today

SambaNovachip vendor

San Jose, CA · San Jose, California, United States · staff plus

custom-asicsoftware-engineercollectivesgpu-kernelsml-platform

posted 1d ago · verified today

Scale AIai startup

San Francisco, CA · San Francisco, CA; Seattle, WA; New York, NY · mid

ml-platformresearch-engineersoftware-engineertraining-frameworksinference-engines

posted 1d ago · verified today

xAIfrontier lab

Palo Alto, CA · Palo Alto, California; Seattle, Washington +1 more · mid

network-fabricsoftware-engineergpu-genericnetwork-engineercollectives

posted 1d ago · verified today

CoreWeaveneocloud

Livingston, NJ · Livingston, NJ / New York, NY +1 more · senior

network-fabricsolutions-architectnetwork-engineernvidiacollectives

posted 1d ago · verified today

CoreWeaveneocloud

Bellevue, WA · Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA +2 more · mid

gpu-genericnvidiasoftware-engineercluster-datacenternetwork-fabric

posted 1d ago · verified today

OpenAIfrontier lab

San Francisco · remote · $250K–$445K · mid

performance-engineergpu-kernelstraining-frameworkscollectivespre-training

posted 10mo ago · verified today

OpenAIfrontier lab

San Francisco · $380K–$555K · mid

collectivessoftware-engineernetwork-fabricnvidiapre-training

posted 28mo ago · verified today

OpenAIfrontier lab

San Francisco · remote · $266K–$445K · mid

custom-asicnetwork-fabricsoftware-engineercluster-datacenterperformance-engineer

posted 10w ago · verified today

OpenAIfrontier lab

San Francisco · $310K–$460K

ml-platformsoftware-engineercluster-datacenterpre-trainingreliability-sre

posted 6mo ago · verified today

OpenAIfrontier lab

London, UK · New York City +2 more · remote · $230K–$405K

gpu-genericcluster-datacentersoftware-engineerscheduling-orchestrationcollectives

posted 3d ago · verified today

OpenAIfrontier lab

San Francisco · Seattle · remote · $293K–$385K · mid

collectivesperformance-engineerdistributed-inferencegpu-genericinference

posted 4mo ago · verified today

Anthropicfrontier lab

San Francisco, CA · San Francisco, CA | New York City, NY | Seattle, WA

gpu-kernelsperformance-engineergpu-genericnvidiadistributed-inference

posted 1d ago · verified today

Anthropicfrontier lab

New York City, NY · Remote-Friendly US (Travel Required) +3 more · senior

gpu-kernelsperformance-engineerinferencecollectivesdistributed-inference

posted 1d ago · verified today

Anthropicfrontier lab

San Francisco, CA · San Francisco, CA | New York City, NY | Seattle, WA · senior

network-fabricsoftware-engineernetwork-engineercollectivesnvidia

posted 1d ago · verified today

Anthropicfrontier lab

New York City, NY · San Francisco, CA +2 more · staff plus

reliability-sresregpu-genericinferenceinference-engines

posted 1d ago · verified today

Nebiusneocloud

Amsterdam, Netherlands; Remote - Europe · United States · remote · senior

cluster-datacenterperformance-engineergpu-genericnetwork-fabriccollectives

posted 1d ago · verified today

Nebiusneocloud

Remote - United States · United States · remote · manager

gpu-genericperformance-engineergpu-kernelscollectivesnetwork-fabric

posted 1d ago · verified today

Crusoeneocloud

San Francisco, CA - US · Sunnyvale, CA - US · staff plus

cluster-datacenterdatacenter-engineergpu-genericnvidiareliability-sre

posted 3mo ago · verified today

Crusoeneocloud

Bellevue, WA - US · San Francisco, CA - US +1 more · staff plus

network-engineernetwork-fabricreliability-sresrecluster-datacenter

posted 4d ago · verified today

Crusoeneocloud

San Francisco, CA - US · staff plus

network-fabricsoftware-engineercluster-datacentergpu-genericml-platform

posted 12d ago · verified today

Crusoeneocloud

San Francisco, CA - US · Sunnyvale, CA - US · senior

cluster-datacentercollectivesgpu-genericnetwork-fabricnvidia

posted 3mo ago · verified today

Crusoeneocloud

San Francisco, CA - US · Sunnyvale, CA - US · staff plus

reliability-sresregpu-genericscheduling-orchestrationcluster-datacenter

posted 6mo ago · verified today

Lambdaneocloud

Bellevue Office · San Francisco Office (Fremont St) +1 more · remote · $314K–$465K · staff plus

scheduling-orchestrationsoftware-engineernvidiagpu-genericml-platform

posted 11d ago · verified today

Lambdaneocloud

Remote, USA · remote · $122K–$162K · mid

cluster-datacentergpu-genericnvidiareliability-sresre

posted 17d ago · verified today