软件工程师,硬件健康
查看雇主原标题
Software Engineer, Hardware HealthOpenAI · San Francisco · $250k – $445k
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 57/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
- 已披露薪资
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 已披露薪资+15
- 稀有职位+1
- 公司来源健康度+8
该职位未包含:远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
机器翻译关于团队 硬件健康与可观测性团队负责 OpenAI 全球计算集群的端到端健康生命周期。 我们的使命是通过可靠的健康信号、自动化修复和可扩展的运维工具,在加速器厂商、代际、云服务商和区域之间最大化健康、可用的算力。 我们构建的系统用于观测、检测、修复和验证 GPU、CPU、网络和平台基础设施中的硬件问题,使前沿模型训练和推理工作负载能够在大规模环境下可靠运行。我们是 OAI 生产和研究工作负载成功的最后一道防线。
关于该职位 在硬件健康与可观测性团队,你将构建关键基础设施,使 OpenAI 最大的计算集群保持健康并大规模运行。 即使少量不健康的系统也可能影响大规模训练和推理工作负载。该团队专注于最大限度减少停机时间、提升集群效率,并确保计算资源持续可供研究人员和产品团队使用。 该团队的工程师端到端负责问题,从定义健康信号和调试故障,到构建可在全球数百万块 GPU 上运行的自动化修复系统。 在这个职位中,你将: • 定义并维护跨 GPU、CPU、网络和平台基础设施的健康信号。
• 构建并演进健康检查,以大规模检测、修复和验证故障。
• 确保关键健康检查以最低延迟执行,以最大化工作负载正常运行时间。
• 调查大规模计算环境中的硬件故障和系统级问题。
• 负责节点生命周期工作流,包括 drain、quarantine、repair、RMA 和 return-to-service 流程。
• 构建自动化和工具,使全球集群管理能够在最少人工干预下进行。
• 与工作负载、可靠性和供应商团队合作,将健康信号集成到训练和推理系统中。
如果你具备以下条件,你可能会在这个职位中如鱼得水: • 7 年以上软件或基础设施工程行业经验。
• 精通 Python 和 shell 脚本。
• 具备构建大规模
岗位职责
在硬件健康与可观测性团队,你将构建关键基础设施,使 OpenAI 最大的计算集群保持健康并大规模运行。 即使少量不健康的系统也可能影响大规模训练和推理工作负载。该团队专注于最大限度减少停机时间、提升集群效率,并确保计算资源持续可供研究人员和产品团队使用。 该团队的工程师端到端负责问题,从定义健康信号和调试故障,到构建可在全球数百万块 GPU 上运行的自动化修复系统。 在这个职位中,你将: • 定义并维护跨 GPU、CPU、网络和平台基础设施的健康信号。
• 构建并演进健康检查,以大规模检测、修复和验证故障。
• 确保关键健康检查以最低延迟执行,以最大化工作负载正常运行时间。
• 调查大规模计算环境中的硬件故障和系统级问题。
• 负责节点生命周期工作流,包括 drain、quarantine、repair、RMA 和 return-to-service 流程。
• 构建自动化和工具,使全球集群管理能够在最少人工干预下进行。
• 与工作负载、可靠性和供应商团队合作,将健康信号集成到训练和推理系统中。
如果你具备以下条件,你可能会在这个职位中如鱼得水: • 7 年以上软件或基础设施工程行业经验。
• 精通 Python 和 shell 脚本。
• 具备构建大规模分布式系统或基础设施平台的经验。
• 能够熟练使用 SQL、PromQL 或类似工具深入分析嘈杂的运维数据。
• 具备构建可复现分析和运维工具的经验。
• 具备出色的系统调试和运维直觉,并拥有主人翁心态。
如果你还具备以下条件,则更佳: • 具备底层硬件系统和 Linux 工具经验(例如 PCIe、InfiniBand、RoCE、网络、电源管理、内核性能调优、FW/SW 调试)。
• 具备运维或调试大规模 GPU 或加速器集群的经验。
• 在网络运维、可观测性或系统遥测方面具备专业知识。
• 具备自动化修复系统或集群生命周期管理经验。
• 具备在分布式计算环境中提升可靠性、利用率或工作负载正常运行时间的经验。
关于 OpenAI OpenAI 是一家 AI 研究和部署公司,致力于确保通用人工智能造福全人类。我们不断拓展 AI 系统能力的边界,并寻求通过我们的产品将其安全地部署到世界各地。AI 是一种极其强大的工具,其创造必须以安全和人类需求为核心;为实现我们的使命,我们必须包容并重视构成人类完整光谱的众多不同视角、声音和经历。 我们是一家提供平等机会的雇主,我们不会基于种族、宗教、肤色、国籍、性别、性取向、年龄、退伍军人身份、残疾、遗传信息或其他适用的受法律保护特征进行歧视。 如需更多信息,请参阅 OpenAI 的平权行动和平等就业机会政策声明。 申请人的背景调查将根据适用法律进行,对于美国候选人,有逮捕或定罪记录的合格申请人将根据相关法律被考虑录用,包括《旧金山公平机会条例》、《洛杉矶县雇主公平机会条例》和《加州公平机会法》。对于未建制洛杉矶县的员工:我们合理认为,犯罪历史可能与以下工作职责存在直接、不利和负面的关系,可能导致有条件录用通知被撤回:保护委托给你的计算机硬件免遭盗窃、丢失或损坏;在雇佣终止或任务结束时归还你持有的所有计算机硬件(包括其中包含的数据);以及维护专有、机密和非公开信息的保密性。此外,工作职责要求访问安全和受保护的信息技术系统以及相关的数据安全义务。 如要通知 OpenAI 你认为该职位发布不合规,请通过此表单提交报告。与职位发布合规无关的询问将不会得到回复。 我们致力于为残障申请人提供合理便利,可通过此链接提出请求。 OpenAI 全球申请人隐私政策 在 OpenAI,我们相信人工智能有潜力帮助人们解决巨大的全球挑战,我们希望 AI 带来的益处能够被广泛共享。加入我们,共同塑造技术的未来。
以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。
查看雇主原文
职位描述
About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: • Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
• Build and evolve health checks that detect, remediate, and verify failures at scale.
• Ensure critical health checks execute with minimal latency to maximize workload uptime.
• Investigate hardware failures and system-level issues across large-scale compute environments.
• Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes.
• Build automation and tooling that enables global cluster management with minimal manual intervention.
• Partner with workload, reliability, and provider teams to integrate health signals into training and inference systems.
You might thrive in this role if you have: • 7+ years of industry experience in software or infrastructure engineering.
• Strong proficiency with Python and shell scripting.
• Experience building large-sc
岗位职责
On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: • Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
• Build and evolve health checks that detect, remediate, and verify failures at scale.
• Ensure critical health checks execute with minimal latency to maximize workload uptime.
• Investigate hardware failures and system-level issues across large-scale compute environments.
• Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes.
• Build automation and tooling that enables global cluster management with minimal manual intervention.
• Partner with workload, reliability, and provider teams to integrate health signals into training and inference systems.
You might thrive in this role if you have: • 7+ years of industry experience in software or infrastructure engineering.
• Strong proficiency with Python and shell scripting.
• Experience building large-scale distributed systems or infrastructure platforms.
• Comfort digging into noisy operational data using SQL, PromQL, or similar tooling.
• Experience building reproducible analyses and operational tooling.
• Strong systems debugging and operational instincts with an ownership mindset.
Bonus if you have: • Experience with low-level hardware systems and Linux tooling (e.g. PCIe, InfiniBand, RoCE, networking, power management, kernel performance tuning, FW/SW debugging).
• Experience operating or debugging large-scale GPU or accelerator clusters.
• Expertise in network operations, observability, or systems telemetry.
• Experience with automated remediation systems or fleet lifecycle management.
• Experience improving reliability, utilization, or workload uptime in distributed compute environments.
About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement . Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link . OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.