跳到主要内容
OOfficialJobs
菜单
官方来源官方来源职位

高级软件工程师,AI Runtime

机器翻译
查看雇主原标题Senior Software Engineer, AI Runtime

Databricks · Mountain View, California; San Francisco, California

职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。

为什么值得关注?

发现指数 42/100,仅依据与该职位一起存储的证据计算。

42/100 发现指数
  • 新的雇主官方职位

分数构成

  • 时效性 (随职位发布时间变化)+18
  • 雇主官方来源+15
  • 稀有职位+1
  • 公司来源健康度+8

该职位未包含:已披露薪资、远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。

这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。

职位描述

机器翻译

P-1428

在 Databricks,我们热衷于帮助数据团队解决世界上最棘手的问题——从让下一代交通方式成为现实,到加速医学突破的研发。我们通过构建和运行全球最佳的数据与 AI 基础设施平台来实现这一目标,让我们的客户能够利用深度数据洞察来改进其业务。

训练和定制最先进的 AI 模型是计算领域要求最苛刻的工作负载之一,而它正是 Databricks Mosaic AI 使命的核心。AI Runtime(AIR)是我们用于大规模 GPU 训练和微调的托管平台。它让客户能够按需访问最新加速器集群,并提供无服务器体验,隐藏了预置、调度和编排多节点作业的复杂性,同时具备韧性,可在数千个 GPU 上持续训练数天或数周。AIR 为全方位的定制训练提供支持,从微调开源模型到预训练前沿规模的基础模型,服务于世界上一些最先进的 AI 团队。

作为 AI Runtime 的高级软件工程师,您将在构建和扩展系统方面发挥关键作用,使大规模训练变得快速、可靠且轻松。您将推动托管 GPU 训练栈的架构和演进,涵盖调度与容量、分布式训练性能、容错,以及大规模启动作业和运维作业的开发者体验。除了对核心系统的亲身贡献外,您还将帮助塑造 AIR 的技术方向,指导其他工程师,与产品、研究和平台团队合作,并为扩大 Databricks 定制训练的技术和业务影响力的各项举措做出贡献。

您将产生的影响:

• 推动 AIR 托管 GPU 训练平台的架构和演进,在跨越数千个加速器的集群上交付可扩展、高吞吐且具有韧性的训练。

• 解决大规模训练中最困难的问题,包括多节点编排、分布式并行策略、GPU 调度和动态路由、高吞吐数据加载,以及针对超长运行作业的检查点和恢复。

• 提升 GPU 效率和训练性能,提高利用率(如模型 FLOPs 利用率和端到端吞吐量),并降低跨不同模型架构和硬件代次的每次训练运行成本。

• 构建韧性和可观测性基础,保持多节点作业健康运行,检测并从硬件和软件故障中恢复,同时将对客户的干扰降至最低。

• 与产品、研究和平台团队合作,塑造 API、CLI 和开发者体验,使启动、监控和调试生产训练作业变得轻松。

• 主导端到端工程工作,从设计到生产上线,对性能、正确性和可靠性保持高标准。

• 为 AIR 背后的核心系统做出直接、高影响力的贡献,并随着集群增长,帮助支持最新加速器和新区域。

• 倡导工程卓越,通过设计评审和技术讨论指导其他工程师,并为 Databricks 在 AI 训练基础设施方面的技术方向做出贡献。

我们寻找的人才:

• 5 年以上构建和运维大规模分布式系统的经验,并具备 GPU 训练基础设施、高性能计算或 ML 系统方面的经验。

• 具备分布式训练框架(如 PyTorch、FSDP、DeepSpeed 或 Megatron)以及用于训练大模型的并行策略(数据并行、张量并行、流水线并行和序列并行)方面的经验。

• 对训练韧性模式有深入理解,包括检查点、故障检测,以及针对长时间运行的多节点作业的自动恢复。

• 扎实掌握 GPU 性能基础知识,包括加速器架构、高速互连(如 NVLink 和 InfiniBand 或 RoCE)、集合通信,以及决定训练吞吐量和利用率的瓶颈。

• 具备在云中构建和运维托管型多租户平台产品的经验,并对可用性、性能和可靠性有明确的 SLA 和 SLO。

• 在算法、数据结构和系统设计方面有扎实基础,并应用于性能敏感的大规模分布式系统。

• 具备交付技术复杂、高影响力举措并创造明确客户或业务价值的可靠能力。

• 出色的沟通能力,以及在快节奏环境中跨产品、研究和基础设施团队协作的能力。

• 以客户为中心的心态,能够将实现细节与产品目标对齐,并热衷于指导工程师和培养技术卓越。

• 计算机科学或相关领域的学士学位(硕士或博士优先)。

福利待遇

在 Databricks,我们致力于提供全面的福利和待遇,以满足所有员工的需求。如需了解您所在地区所提供福利的具体详情,请点击此处。

我们对多元与包容的承诺

在 Databricks,我们致力于营造多元和包容的文化,让每个人都能脱颖而出。我们非常重视确保我们的招聘实践具有包容性,并符合平等就业机会标准。在 Databricks 寻求就业机会的个人,不会因年龄、肤色、残疾、族裔、家庭或婚姻状况、性别认同或表达、语言、国籍、身体和心理能力、政治派别、种族、宗教、性取向、社会经济状况、退伍军人身份及其他受保护特征而受到区别对待。

合规

如果履行工作职责需要访问受出口管制的技术或源代码,雇主可自行决定是否为此类职位申请美国政府许可证,且雇主可能仅基于此原因而拒绝继续推进某位申请人的流程。

薪资

Databricks 致力于公平和公正的薪酬实践。该职位的薪资范围列于下方,代表非佣金制职位的预期薪资范围或佣金制职位的目标收入。实际薪酬方案取决于每位候选人独有的若干因素,包括但不限于与工作相关的技能、经验深度、相关认证和培训,以及具体工作地点。基于上述因素,Databricks 预计将使用该范围的全部宽度。该职位的总薪酬方案还可能包括年度绩效奖金、股权以及上述福利的资格。如需了解您所在地点属于哪个范围的更多信息,请访问我们的页面此处。

当地薪资范围 $160,000 — $225,000 USD

关于 Databricks

Databricks 是一家数据与 AI 公司。全球超过 20,000 家组织——包括 adidas、AT&T、Bayer、Block、Mastercard、Rivian、Unilever,以及 70% 的《财富》500 强企业——依赖 Databricks Data + AI Platform 来构建和扩展数据与 AI 应用、分析和智能体。Databricks 总部位于旧金山,在全球拥有 30 多个办事处,提供统一平台,包括 Genie、Lakebase、Agent Bricks、Lakeflow、Lakehouse 和 Unity Catalog。如需了解更多信息,请在 LinkedIn、X、YouTube 和 Instagram 上关注 Databricks。

以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。

查看雇主原文

职位描述

P-1428

At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business.

Training and customizing state-of-the-art AI models is one of the most demanding workloads in computing, and it sits at the heart of Databricks' Mosaic AI mission. AI Runtime (AIR) is our managed platform for large-scale GPU training and fine-tuning. It gives customers on-demand access to fleets of the latest accelerators and a serverless experience that hides the complexity of provisioning, scheduling, and orchestrating multi-node jobs, with the resilience to keep training running for days or weeks across thousands of GPUs. AIR powers the full spectrum of custom training, from fine-tuning open models to pre-training frontier-scale foundation models, for some of the most sophisticated AI teams in the world.

As a Senior Software Engineer for AI Runtime, you will play a critical role in building and scaling the systems that make large-scale training fast, reliable, and effortless. You will drive the architecture and evolution of the managed GPU training stack, spanning scheduling and capacity, distributed training performance, fault tolerance, and the developer experience of launching and operating jobs at scale. Beyond hands-on contributions to core systems, you will help shape the technical direction for AIR, mentor other engineers, partner across product, research, and platform teams, and contribute to the initiatives that expand the technical and business impact of custom training at Databricks.

The impact you will have:

• Drive the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets that span thousands of accelerators.

• Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs.

• Push GPU efficiency and training performance, raising utilization (such as model FLOPs utilization and end-to-end throughput) and lowering cost per training run across diverse model architectures and hardware generations.

• Build the resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal disruption to customers.

• Partner with product, research, and platform teams to shape the APIs, CLI, and developer experience that make it easy to launch, monitor, and debug production training jobs.

• Lead end-to-end engineering efforts, from design through production rollout, holding a high bar for performance, correctness, and reliability.

• Make direct, high-impact contributions to the core systems behind AIR, and help bring up support for the latest accelerators and new regions as the fleet grows.

• Champion engineering excellence, mentor other engineers through design reviews and technical discussions, and contribute to Databricks' technical direction in AI training infrastructure.

What we look for:

• 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, high-performance computing, or ML systems.

• Experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models.

• Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.

• Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (such as NVLink and InfiniBand or RoCE), collective communication, and the bottlenecks that govern training throughput and utilization.

• Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability.

• Strong foundation in algorithms, data structures, and system design as applied to performance-sensitive, large-scale distributed systems.

• Proven ability to deliver technically complex, high-impact initiatives that create clear customer or business value.

• Strong communication skills and the ability to collaborate across product, research, and infrastructure teams in a fast-moving environment.

• Customer-focused mindset with the ability to align implementation details with product goals, and a passion for mentoring engineers and fostering technical excellence.

• BS in Computer Science or a related field (MS or PhD preferred).

福利待遇

At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here .

Our Commitment to Diversity and Inclusion

At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.

Compliance

If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.

薪资

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here .

Local Pay Range $160,000 — $225,000 USD

About Databricks

Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn , X , YouTube , and Instagram .

Databricks 的更多职位

公司主页
官方来源最新
Bellevue, Washington; Mountain View, California; New York City, New York; San Francisco, California; Seattle, Washington; Washington, D.C.全职From 1
英文原文

GAQ427R191 Databricks is seeking a Senior Product Counsel to help the Databricks legal team provide cutting-edge, practical legal advice to Databricks’ engineering and product management teams as w…

官方来源职位
首次发现于8小时前
已核实8小时前
官方来源最新
Tokyo, 日本全职未披露薪资
英文原文

Location: Tokyo, Japan Req ID: FEQ427R141 Reporting to: Manager, Field Engineering Language requirements: Business-professional Japanese required; business-level English preferred Mission As…

官方来源职位
首次发现于8小时前
已核实8小时前
官方来源最新
Paris, 法国全职未披露薪资
英文原文

FEQ427R550 Solutions Architect - Portuguese speaking. Location: Madrid or Paris, or within a commutable distance The Role As a Solutions Architect, you will lead the technical strategy fo…

官方来源职位
首次发现于8小时前
已核实8小时前
官方来源最新
Madrid全职未披露薪资
英文原文

FEQ427R550 Solutions Architect - Portuguese speaking. Location: Madrid or Paris, or within a commutable distance The Role As a Solutions Architect, you will lead the technical strategy fo…

官方来源职位
首次发现于8小时前
已核实8小时前

其他公司的相似职位

搜索这类职位
官方来源最新
Hybrid - San Francisco, New York City混合办公全职$208k – $312k
英文原文

About Vercel: Vercel is the agentic infrastructure company. We free people and agents to ship what’s next. For more than a decade, Vercel has shaped how the web is built. As the team behind Next…

未出现在监控的职位板上
首次发现于2小时前
已核实2小时前

Staff Software Engineer, Event logging原文

Airbnb · Software Engineering

官方来源最新
美国全职未披露薪资
英文原文

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every c…

未出现在监控的职位板上
首次发现于2小时前
已核实2小时前

Senior Staff Software Engineer, Trust原文

Airbnb · Software Engineering

官方来源最新
Remote - US远程全职未披露薪资
英文原文

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every c…

未出现在监控的职位板上
首次发现于2小时前
已核实2小时前
官方来源最新
美国全职未披露薪资
英文原文

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every c…

未出现在监控的职位板上
首次发现于2小时前
已核实2小时前