软件工程师,计算基础系统
查看雇主原标题
Software Engineer, Compute Foundations SystemsOpenAI · San Francisco · $230k – $490k
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 57/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
- 已披露薪资
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 已披露薪资+15
- 稀有职位+1
- 公司来源健康度+8
该职位未包含:远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
机器翻译关于团队 Frontier Systems Foundations 隶属于 OpenAI 的 Compute Foundations,负责构建系统软件基础,将新的计算基础设施转化为可靠、可用的容量,用于前沿模型训练。 我们的使命是让一些全球最大的 GPU 集群可靠地服务于前沿训练。我们让新平台和集群上线,安全维护已安装的机群,并与硬件、基础设施和研究团队合作,解决阻碍作业运行的系统级问题。 这意味着构建和维护最贴近机器的软件:Linux 和 Ubuntu 操作系统镜像、内核和模块、驱动、软件包和仓库、磁盘和启动配置、固件集成、置备以及系统级验证。我们让这些组件在异构机群中可复现、兼容且可安全运行。
关于该职位 我们正在寻找具备深厚 Linux 和主机系统经验的系统软件工程师,来构建、验证和维护 OpenAI 前沿计算机群的操作系统基础。相关背景包括内核和模块开发、Linux 发行版或镜像工程、软件包管理、固件和驱动集成、磁盘和启动,以及裸金属置备。 你将与硬件工程师、供应商和基础设施团队紧密合作,让新平台上线、集成系统组件,并调试固件、磁盘、启动、操作系统、内核、驱动以及工作负载交互中的故障。你的工作将直接影响新容量多快能够投入使用,以及大型 GPU 机群运行得有多可靠。 你应该能够自如地编写和维护生产级系统软件与自动化,但我们并不期望你精通每一层。这是一个深入攻克具有挑战性的系统问题的机会,同时构建支撑下一代前沿模型的镜像、软件包、验证和恢复路径。
在这个职位中,你将: • 为大型 GPU 集群构建和维护 Linux 主机软件栈,包括 Ubuntu 和操作系统镜像、内核配置和模块、驱动、软件包、磁盘和存储配置,以及机器配置。
• 为异构裸金属和云计算机群设计可复现的操作系统镜像构建、软件包和仓库工作流,以及系统配置。
• 在新旧硬件平台上集成、测试和验证内核、模块、驱动、软件包和固件;为系统变更构建安全的金丝雀、回滚和恢复路径。
• 让新硬件平台和计算 SKU 上线,与硬件工程师和供应商合作解决固件、驱动、操作系统和兼容性问题。
• 调试启动和置备、固件、磁盘、内核和驱动,以及工作负载交互中的复杂系统故障;将反复出现的故障模式转化为持久的修复、测试和自动化。
• 通过消除手动逐主机 h,提升置备、重装、维护和恢复的正确性
岗位职责
我们正在寻找具有深厚 Linux 和主机系统经验的系统软件工程师,来构建、验证和维护 OpenAI 前沿计算集群的操作系统基础。相关背景包括内核和模块开发、Linux 发行版或镜像工程、软件包管理、固件和驱动集成、磁盘与启动,以及裸金属配置。 你将与硬件工程师、供应商和基础设施团队紧密合作,引入新平台、集成系统组件,并调试固件、磁盘、启动、操作系统、内核、驱动以及工作负载交互中的故障。你的工作将直接影响新容量能以多快速度投入使用,以及大型 GPU 集群能以多高可靠性运行。 你应该能够自如地编写和维护生产级系统软件与自动化,但我们并不期望你精通每一层。这是一个深入攻克具有挑战性的系统问题的机会,同时构建支撑下一代前沿模型的镜像、软件包、验证和恢复路径。 在这个职位中,你将: • 为大型 GPU 集群构建和维护 Linux 主机软件栈,包括 Ubuntu 和操作系统镜像、内核配置和模块、驱动、软件包、磁盘和存储配置,以及机器配置。
• 为异构裸金属和云计算集群设计可复现的操作系统镜像构建、软件包和仓库工作流,以及系统配置。
• 在新旧硬件平台上集成、测试和验证内核、模块、驱动、软件包和固件;为系统变更构建安全的金丝雀、回滚和恢复路径。
• 引入新的硬件平台和计算 SKU,与硬件工程师和供应商合作解决固件、驱动、操作系统和兼容性问题。
• 调试跨启动和配置、固件、磁盘、内核和驱动,以及工作负载交互的复杂系统故障;将反复出现的故障模式转化为持久修复、测试和自动化。
• 通过消除逐台主机的人工干预,并使系统行为在集群规模下可预测,提升配置、重装、维护和恢复的正确性。
• 构建有针对性的系统工具和诊断能力,使主机软件栈在已部署集群中更易于验证、排查和运维。
如果你符合以下条件,你可能会在这个职位中如鱼得水: • 在构建、集成或运维生产 Linux 系统方面拥有丰富经验,尤其是 Ubuntu、Debian 或其他主流 Linux 发行版。
• 在以下一个或多个领域拥有深厚经验:Linux 内核、模块或设备驱动;Linux 发行版、软件包、仓库或操作系统镜像工程;启动、配置、磁盘或机器配置;固件和驱动集成、验证或推广。
• 能够使用合适的系统语言、脚本和 Linux 工具编写、调试和维护生产级系统软件与自动化。
• 理解如何构建、测试、打包、验证并安全交付系统级变更,包括兼容性测试、分阶段推广、回滚和恢复。
• 能够系统性地调试跨越硬件、固件、启动、操作系统、内核和驱动边界的故障,并与硬件和软件合作伙伴高效协作。
• 乐于调查困难的系统行为,识别底层故障模式,并构建务实的修复方案,使生产系统更加可靠。
如果你具备以下条件,将获得加分: • 构建过裸金属配置、机器引导、操作系统重新配置或系统生命周期工具。
• 使用过 GPU 系统、AI 基础设施、HPC 集群、加速器或其他异构且对性能敏感的计算环境。
• 有引入新服务器平台的经验,或直接与硬件和软件供应商合作处理固件、驱动或操作系统兼容性问题。
• 为 Linux、内核子系统或模块、发行版、软件包生态、驱动、固件工具或其他开源系统软件做出过贡献。
• 在大型生产计算集群中交付或支持过系统软件变更。
关于 OpenAI OpenAI 是一家 AI 研究与部署公司,致力于确保通用人工智能造福全人类。我们推动 AI 系统能力边界,并寻求通过我们的产品将其安全部署到世界。AI 是一种极其强大的工具,必须以安全和人类需求为核心来创造;为实现我们的使命,我们必须包容并重视构成人类完整光谱的众多不同视角、声音和经验。 我们是提供平等机会的雇主,我们不会基于种族、宗教、肤色、国籍、性别、性取向、年龄、退伍军人身份、残疾、遗传信息或其他适用的受法律保护特征进行歧视。 如需更多信息,请参阅 OpenAI 的平权行动和 equal employment opportunity 政策声明。 对申请人的背景调查将依据适用法律进行,对于美国候选人,有逮捕或定罪记录的合格申请人将依据这些法律获得就业考虑,包括《旧金山公平机会条例》、《洛杉矶县雇主公平机会条例》和《加州公平机会法》。对于未建制洛杉矶县的员工:我们合理认为犯罪历史可能与以下工作职责存在直接、不利和负面的关系,可能导致撤回有条件录用通知:保护委托给你的计算机硬件免遭盗窃、丢失或损坏;在雇佣终止或任务结束时归还你持有的所有计算机硬件(包括其中包含的数据);以及维护专有、机密和非公开信息的保密性。此外,工作职责要求访问安全和受保护的信息技术系统以及相关的数据安全义务。 如需通知 OpenAI 你认为该职位发布不合规,请通过此表单提交报告。与职位发布合规无关的询问将不会得到回复。 我们致力于为残障申请人提供合理便利,可通过此链接提出请求。 OpenAI 全球申请人隐私政策 在 OpenAI,我们相信人工智能有潜力帮助人们解决巨大的全球挑战,我们希望 AI 带来的益处能够被广泛共享。加入我们,共同塑造技术的未来。
以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。
查看雇主原文
职位描述
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models. In this role, you will: • Build and maintain the Linux host software stack for large GPU clusters, including Ubuntu and OS images, kernel configuration and modules, drivers, packages, disks and storage configuration, and machine configuration.
• Design reproducible OS-image builds, package and repository workflows, and system configuration for heterogeneous bare-metal and cloud compute fleets.
• Integrate, test, and qualify kernels, modules, drivers, packages, and firmware across new and existing hardware platforms; build safe canary, rollback, and recovery paths for system changes.
• Bring up new hardware platforms and compute SKUs, working with hardware engineers and vendors to resolve firmware, driver, operating-system, and compatibility issues.
• Debug complex system failures across boot and provisioning, firmware, disks, kernels and drivers, and workload interactions; turn recurring failure modes into durable fixes, tests, and automation.
• Improve provisioning, repave, maintenance, and recovery correctness by eliminating manual host-by-h
岗位职责
We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models. In this role, you will: • Build and maintain the Linux host software stack for large GPU clusters, including Ubuntu and OS images, kernel configuration and modules, drivers, packages, disks and storage configuration, and machine configuration.
• Design reproducible OS-image builds, package and repository workflows, and system configuration for heterogeneous bare-metal and cloud compute fleets.
• Integrate, test, and qualify kernels, modules, drivers, packages, and firmware across new and existing hardware platforms; build safe canary, rollback, and recovery paths for system changes.
• Bring up new hardware platforms and compute SKUs, working with hardware engineers and vendors to resolve firmware, driver, operating-system, and compatibility issues.
• Debug complex system failures across boot and provisioning, firmware, disks, kernels and drivers, and workload interactions; turn recurring failure modes into durable fixes, tests, and automation.
• Improve provisioning, repave, maintenance, and recovery correctness by eliminating manual host-by-host intervention and making system behavior predictable at fleet scale.
• Build focused systems tooling and diagnostics that make the host software stack easier to validate, troubleshoot, and operate across the installed fleet.
You might thrive in this role if you: • Have significant experience building, integrating, or operating production Linux systems, particularly Ubuntu, Debian, or another major Linux distribution.
• Have deep experience in one or more of: Linux kernels, modules, or device drivers; Linux distribution, package, repository, or OS-image engineering; boot, provisioning, disks, or machine configuration; firmware and driver integration, qualification, or rollout.
• Can write, debug, and maintain production-quality systems software and automation using appropriate systems languages, scripting, and Linux tooling.
• Understand how to build, test, package, qualify, and safely deliver system-level changes, including compatibility testing, staged rollout, rollback, and recovery.
• Can systematically debug failures that cross hardware, firmware, boot, operating-system, kernel, and driver boundaries, and collaborate effectively with hardware and software partners.
• Enjoy investigating difficult system behavior, identifying underlying failure modes, and building pragmatic fixes that make production systems more reliable.
Bonus points if you: • Have built bare-metal provisioning, machine-bootstrap, OS-reprovisioning, or system-lifecycle tooling.
• Have worked with GPU systems, AI infrastructure, HPC clusters, accelerators, or other heterogeneous and performance-sensitive compute environments.
• Have experience bringing up new server platforms or working directly with hardware and software vendors on firmware, driver, or operating-system compatibility issues.
• Have contributed to Linux, kernel subsystems or modules, distributions, package ecosystems, drivers, firmware tooling, or other open-source systems software.
• Have delivered or supported system-software changes across large production compute fleets.
About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement . Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link . OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.