跳到主要内容
OOfficialJobs
菜单
官方来源官方来源职位

软件工程师,GPU 基础设施 - HPC

机器翻译
查看雇主原标题Software Engineer, GPU Infrastructure - HPC

OpenAI · San Francisco; New York City · $230k – $490k

职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。

为什么值得关注?

发现指数 57/100,仅依据与该职位一起存储的证据计算。

57/100 发现指数
  • 新的雇主官方职位
  • 已披露薪资

分数构成

  • 时效性 (随职位发布时间变化)+18
  • 雇主官方来源+15
  • 已披露薪资+15
  • 稀有职位+1
  • 公司来源健康度+8

该职位未包含:远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。

这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。

职位描述

机器翻译

关于团队 OpenAI 的 Fleet 团队为支撑我们前沿研究和产品开发的计算环境提供支持。我们负责管理涵盖数据中心、GPU、网络等的大规模系统,确保高可用性、高性能和高效率。我们的工作使 OpenAI 的模型能够大规模无缝运行,同时支持内部研究以及 ChatGPT 等外部产品。我们优先考虑安全性、可靠性和负责任的 AI 部署,而非无节制的增长。

关于该职位 作为 Fleet 高性能计算(HPC)团队的软件工程师,你将负责 OpenAI 全部计算集群的可靠性和正常运行时间。最大限度地减少硬件故障是研究训练进展和服务稳定的关键,因为哪怕一次硬件小故障都可能造成重大中断。随着超级计算机规模日益庞大,风险也在不断上升。

身处技术前沿意味着我们往往是规模化排查这些最先进系统的先驱。这是一个独特的机会,可以接触前沿技术并设计创新解决方案,以维护我们超级计算基础设施的健康和效率。

我们的团队为优秀工程师赋能,给予他们高度的自主权和主人翁意识,以及推动变革的能力。该职位需要高度专注于系统层面的全面调查以及自动化解决方案的开发。我们希望找到这样的人:深入钻研问题,尽可能彻底地调查,并构建用于大规模检测和修复的自动化系统。

在这个职位中,你将: • 构建并维护用于配置和管理服务器集群的自动化系统。

• 开发用于监控服务器健康、性能和生命周期事件的工具。

• 与集群、网络和基础设施团队协作。

• 与外部运营商合作,确保高质量水平。

• 识别并修复性能瓶颈和低效问题。

• 持续改进自动化以减少人工工作。

如果你具备以下条件,你可能会在这个职位上如鱼得水: • 管理大规模服务器环境的经验。

• 在构建和运营方面兼具优势

岗位职责

作为 Fleet 高性能计算(HPC)团队的软件工程师,你将负责 OpenAI 全部计算集群的可靠性和正常运行时间。最大限度地减少硬件故障是研究训练进展和服务稳定的关键,因为哪怕一次硬件小故障都可能造成重大中断。随着超级计算机规模日益庞大,风险也在不断上升。

身处技术前沿意味着我们往往是规模化排查这些最先进系统的先驱。这是一个独特的机会,可以接触前沿技术并设计创新解决方案,以维护我们超级计算基础设施的健康和效率。

我们的团队为优秀工程师赋能,给予他们高度的自主权和主人翁意识,以及推动变革的能力。该职位需要高度专注于系统层面的全面调查以及自动化解决方案的开发。我们希望找到这样的人:深入钻研问题,尽可能彻底地调查,并构建用于大规模检测和修复的自动化系统。

在这个职位中,你将: • 构建并维护用于配置和管理服务器集群的自动化系统。

• 开发用于监控服务器健康、性能和生命周期事件的工具。

• 与集群、网络和基础设施团队协作。

• 与外部运营商合作,确保高质量水平。

• 识别并修复性能瓶颈和低效问题。

• 持续改进自动化以减少人工工作。

如果你具备以下条件,你可能会在这个职位上如鱼得水:

任职要求

• 在构建和运营方面兼具优势。

• 精通 Python、Go 或类似语言。

• 扎实的 Linux、网络和服务器硬件知识。

• 能够自如地使用 SQL、PromQL 和 Pandas 或任何其他工具深入分析嘈杂数据。

该职位不要求具备硬件方面的既往专业知识。

加分技能: • 具备硬件组件、协议及相关 Linux 工具(例如 PCIe、Infiniband、网络、电源管理、内核性能调优)的底层细节经验

• 了解硬件管理协议(例如 IPMI、Redfish)。

• 高性能计算(HPC)或分布式系统经验。

• 具备开发、管理或设计硬件的既往经验。

• 熟悉监控工具(例如 Prometheus、Grafana)。

关于 OpenAI OpenAI 是一家 AI 研究和部署公司,致力于确保通用人工智能造福全人类。我们不断拓展 AI 系统能力的边界,并寻求通过我们的产品将其安全地部署到全世界。AI 是一种极其强大的工具,其创建必须以安全和人类需求为核心;为实现我们的使命,我们必须包容并重视构成人类全貌的众多不同视角、声音和经历。 我们是一家提供平等机会的雇主,我们不会基于种族、宗教、肤色、国籍、性别、性取向、年龄、退伍军人身份、残疾、遗传信息或其他适用的受法律保护特征进行歧视。 如需了解更多信息,请参阅 OpenAI 的《平权行动与平等就业机会政策声明》。 对申请人的背景调查将依照适用法律进行,对于有逮捕或定罪记录的合格申请人,将依据相关法律予以就业考虑,包括针对美国候选人的《旧金山公平机会条例》、《洛杉矶县雇主公平机会条例》和《加州公平机会法》。对于未建制洛杉矶县的员工:我们有合理理由认为,犯罪史可能与以下工作职责存在直接、不利和负面的关系,并可能导致撤回有条件录用通知:保护委托给你的计算机硬件免遭盗窃、丢失或损坏;在雇佣终止或任务结束时归还你持有的所有计算机硬件(包括其中包含的数据);以及对专有、机密和非公开信息保密。此外,工作职责要求访问安全且受保护的信息技术系统,并承担相关的数据安全义务。 如需通知 OpenAI 你认为该职位发布不合规,请通过此表单提交报告。与职位发布合规无关的询问将不予回复。 我们致力于为残障申请人提供合理便利,可通过此链接提出请求。 OpenAI 全球申请人隐私政策 在 OpenAI,我们相信人工智能有潜力帮助人们解决巨大的全球性挑战,我们希望 AI 带来的益处能够被广泛共享。加入我们,共同塑造技术的未来。

以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。

查看雇主原文

职位描述

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: • Build and maintain automation systems for provisioning and managing server fleets.

• Develop tools to monitor server health, performance, and lifecycle events.

• Collaborate with clusters, networking, and infrastructure teams.

• Partner with external operators to ensure a high level of quality.

• Identify and fix performance bottlenecks and inefficiencies.

• Continuously improve automation to reduce manual work.

You might thrive in this role if you have: • Experience managing large-scale server environments.

• A balance of strengths in build

岗位职责

As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: • Build and maintain automation systems for provisioning and managing server fleets.

• Develop tools to monitor server health, performance, and lifecycle events.

• Collaborate with clusters, networking, and infrastructure teams.

• Partner with external operators to ensure a high level of quality.

• Identify and fix performance bottlenecks and inefficiencies.

• Continuously improve automation to reduce manual work.

You might thrive in this role if you have:

任职要求

• A balance of strengths in building and operationalizing.

• Proficiency in Python, Go, or similar languages.

• Strong Linux, networking, and server hardware knowledge.

• Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.

Prior hardware expertise is not required for this role. Bonus Skills: • Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)

• Knowledge of hardware management protocols (e.g., IPMI, Redfish).

• High-performance computing (HPC) or distributed systems experience.

• Prior experience developing, managing, or designing hardware.

• Familiarity with monitoring tools (e.g., Prometheus, Grafana).

About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement . Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link . OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

OpenAI 的更多职位

公司主页
官方来源最新
San Francisco远程全职$293k – $325k
英文原文

About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI mode…

未出现在监控的职位板上
首次发现于14小时前
已核实6小时前

模型策略经理

OpenAI · Safety Systems, Model Policy

官方来源最新
San Francisco远程全职$266k – $335k
英文原文

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency.…

未出现在监控的职位板上提及搬迁
首次发现于14小时前
已核实6小时前

应用人工智能工程师

OpenAI · Go To Market, Technical Success

官方来源最新
新加坡远程全职未披露薪资
英文原文

About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, prod…

未出现在监控的职位板上提及搬迁
首次发现于14小时前
已核实6小时前

战略财务,算力

OpenAI · Strategic Finance, Strategic Finance

官方来源最新
San Francisco全职$234k – $260k
英文原文

About the Team The Compute & Infrastructure Strategy team handles strategy and execution of OpenAI’s compute roadmap. This team’s key responsibilities span financial analysis & reporting, capacity…

未出现在监控的职位板上
首次发现于14小时前
已核实6小时前

其他公司的相似职位

搜索这类职位
India - Bangalore全职未披露薪资
英文原文

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Software Engineering Job Detail…

未出现在监控的职位板上
首次发现于6小时前
已核实6小时前

Systems Engineer原文

Cloudflare · Engineering

官方来源最新
混合办公混合办公全职未披露薪资
英文原文

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet propertie…

官方来源职位
首次发现于6小时前
已核实6小时前
官方来源最新
混合办公混合办公全职未披露薪资
英文原文

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet propertie…

官方来源职位
首次发现于6小时前
已核实6小时前

Software Engineer原文

Coinbase · Engineering - Frontend

官方来源最新
Remote - 加拿大远程全职未披露薪资
英文原文

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that w…

官方来源职位
首次发现于6小时前
已核实6小时前