跳到主要内容
OOfficialJobs
菜单
官方来源官方来源职位

软件工程师,数据基础设施 - 研究

机器翻译
查看雇主原标题Software Engineer, Data Infrastructure - Research

OpenAI · San Francisco · $250k – $380k

职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。

为什么值得关注?

发现指数 57/100,仅依据与该职位一起存储的证据计算。

57/100 发现指数
  • 新的雇主官方职位
  • 已披露薪资

分数构成

  • 时效性 (随职位发布时间变化)+18
  • 雇主官方来源+15
  • 已披露薪资+15
  • 稀有职位+1
  • 公司来源健康度+8

该职位未包含:远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。

这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。

职位描述

机器翻译

关于团队 Workload 团队负责设计和运行 OpenAI 的 LLM 训练与推理基础设施,为大规模的前沿模型提供支持。我们的系统统一了研究人员训练和部署模型的方式,抽象掉了跨庞大 GPU/加速器集群在性能、并行和執行方面的复杂性。通过提供这一基础,Workload 团队确保研究人员能够专注于推进模型能力,而我们则负责处理将这些模型变为现实所需的规模、效率和可靠性。

关于该职位 我们正在寻找一位工程师,负责设计和实现为 OpenAI 下一代训练技术栈提供支持的数据集基础设施。你将负责构建标准化的数据集接口、在数千个 GPU 上扩展流水线,并主动测试性能瓶颈。在这一职位中,你将与多模态研究人员以及其他基础设施团队紧密合作,确保数据集统一、高效且易于使用。

在这一职位中,你将: • 设计并维护标准化的数据集 API,包括针对无法放入内存的多模态(MM)数据。

• 构建主动测试和规模验证流水线,以支持 GPU 规模的数据集加载。

• 与团队成员合作,将数据集无缝集成到训练和推理流水线中,确保顺利采用和出色的用户体验。

• 记录并维护数据集接口,使其易于发现、保持一致,并便于其他团队采用。

• 建立保障和验证系统,确保数据集在标准化后保持可复现且不发生改变。

• 调试并解决分布式数据集加载中的性能瓶颈(例如拖慢全局训练的掉队系统)。

• 提供可视化和检查工具,以呈现数据集中的错误、缺陷或瓶颈。

如果你具备以下条件,你可能会在这一职位中如鱼得水: • 拥有扎实的工程基础,并在分布式系统、数据流水线或基础设施方面有经验。

• 拥有构建 API、模块化代码和可扩展抽象的经验

岗位职责

我们正在寻找一位工程师,负责设计和实现为 OpenAI 下一代训练技术栈提供支持的数据集基础设施。你将负责构建标准化的数据集接口、在数千个 GPU 上扩展流水线,并主动测试性能瓶颈。在这一职位中,你将与多模态研究人员以及其他基础设施团队紧密合作,确保数据集统一、高效且易于使用。

在这一职位中,你将: • 设计并维护标准化的数据集 API,包括针对无法放入内存的多模态(MM)数据。

• 构建主动测试和规模验证流水线,以支持 GPU 规模的数据集加载。

• 与团队成员合作,将数据集无缝集成到训练和推理流水线中,确保顺利采用和出色的用户体验。

• 记录并维护数据集接口,使其易于发现、保持一致,并便于其他团队采用。

• 建立保障和验证系统,确保数据集在标准化后保持可复现且不发生改变。

• 调试并解决分布式数据集加载中的性能瓶颈(例如拖慢全局训练的掉队系统)。

• 提供可视化和检查工具,以呈现数据集中的错误、缺陷或瓶颈。

如果你具备以下条件,你可能会在这一职位中如鱼得水: • 拥有扎实的工程基础,并在分布式系统、数据流水线或基础设施方面有经验。

• 拥有构建 API、模块化代码和可扩展抽象的经验,同时认识到抽象最终服务于用户,而 UX 是抽象设计的重要组成部分。

• 能够自如地调试大规模机器集群中的瓶颈。

• 以构建“开箱即用”的基础设施为荣,并从成为可靠性和规模的守护者中获得乐趣。

• 善于协作、谦逊,并乐于负责 ML 技术栈中基础性(即便不那么光鲜)的部分。

如果你还具备以下条件,将获得加分: • 拥有数据数学、概率或分布式数据理论方面的背景知识。

• 曾参与 GPU 规模的分布式系统或面向实时数据的数据集扩展工作

关于 OpenAI OpenAI 是一家 AI 研究与部署公司,致力于确保通用人工智能造福全人类。我们不断拓展 AI 系统能力的边界,并寻求通过我们的产品将其安全地部署到世界各地。AI 是一种极其强大的工具,其创建必须以安全和人类需求为核心;为实现我们的使命,我们必须包容并重视构成人类全貌的众多不同视角、声音和经历。 我们是一家机会均等的雇主,不会基于种族、宗教、肤色、国籍、性别、性取向、年龄、退伍军人身份、残疾、遗传信息或其他适用的受法律保护特征进行歧视。 如需了解更多信息,请参阅 OpenAI 的《平权行动与平等就业机会政策声明》。 对申请人的背景调查将依据适用法律进行,对于美国境内的候选人,有逮捕或定罪记录的合格申请人将依据相关法律获得就业考虑,包括《旧金山公平机会条例》、《洛杉矶县雇主公平机会条例》和《加州公平机会法》。对于未建制洛杉矶县的员工:我们合理认为,犯罪史可能与以下工作职责存在直接、不利和负面的关系,并可能导致有条件录用通知被撤回:保护委托给你的计算机硬件免遭盗窃、丢失或损坏;在雇佣终止或任务结束时归还你持有的所有计算机硬件(包括其中包含的数据);以及维护专有、机密和非公开信息的保密性。此外,工作职责要求访问安全且受保护的信息技术系统,并承担相关的数据安全义务。 如需通知 OpenAI 你认为该职位发布不合规,请通过此表单提交报告。与职位发布合规无关的询问将不予回复。 我们致力于为残障申请人提供合理的便利,可通过此链接提出请求。 OpenAI 全球申请人隐私政策 在 OpenAI,我们相信人工智能有潜力帮助人们解决巨大的全球性挑战,我们希望 AI 带来的益处能够被广泛共享。加入我们,共同塑造技术的未来。

以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。

查看雇主原文

职位描述

About the Team The Workload team is responsible for designing and running OpenAI’s LLM training and inference infrastructure that powers frontier models at massive scale. Our systems unify how researchers train and serve models, abstracting away the complexity of performance, parallelism, and execution across vast GPU/accelerator fleets. By providing this foundation, the Workload team ensures that researchers can focus on advancing model capabilities while we handle the scale, efficiency, and reliability required to bring those models to life. About the Role We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. In this role, you will: • Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory.

• Build proactive testing and scale validation pipelines for dataset loading at GPU scale.

• Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience.

• Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt.

• Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized.

• Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training).

• Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets.

You might thrive in this role if you: • Have strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.

• Have experience building APIs, modular code, and scalable abstraction

岗位职责

We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. In this role, you will: • Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory.

• Build proactive testing and scale validation pipelines for dataset loading at GPU scale.

• Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience.

• Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt.

• Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized.

• Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training).

• Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets.

You might thrive in this role if you: • Have strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.

• Have experience building APIs, modular code, and scalable abstractions, while recognizing that abstractions ultimately serve the users and UX is an important part of the abstractions design.

• Are comfortable debugging bottlenecks across large fleets of machines.

• Take pride in building infrastructure that “just works,” and find joy in being the guardian of reliability and scale.

• Are collaborative, humble, and excited to own a foundational (if not glamorous) part of the ML stack.

Bonus points if you: • Have background knowledge in data math, probability, or distributed data theory.

• Have worked with GPU-scale distributed systems or dataset scaling for real-time data

About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement . Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link . OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

OpenAI 的更多职位

公司主页
官方来源最新
San Francisco远程全职$293k – $325k
英文原文

About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI mode…

未出现在监控的职位板上
首次发现于14小时前
已核实6小时前

模型策略经理

OpenAI · Safety Systems, Model Policy

官方来源最新
San Francisco远程全职$266k – $335k
英文原文

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency.…

未出现在监控的职位板上提及搬迁
首次发现于14小时前
已核实6小时前

应用人工智能工程师

OpenAI · Go To Market, Technical Success

官方来源最新
新加坡远程全职未披露薪资
英文原文

About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, prod…

未出现在监控的职位板上提及搬迁
首次发现于14小时前
已核实6小时前

战略财务,算力

OpenAI · Strategic Finance, Strategic Finance

官方来源最新
San Francisco全职$234k – $260k
英文原文

About the Team The Compute & Infrastructure Strategy team handles strategy and execution of OpenAI’s compute roadmap. This team’s key responsibilities span financial analysis & reporting, capacity…

未出现在监控的职位板上
首次发现于14小时前
已核实6小时前

其他公司的相似职位

搜索这类职位

Senior Systems Engineer原文

Graphcore · DC Engineering*

官方来源最新
Austin, Texas, 美国全职salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits inc
英文原文

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation o…

官方来源职位
首次发现于刚刚
已核实刚刚
官方来源最新
Sunnyvale全职$21k – $290k
英文原文

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex…

官方来源职位
首次发现于刚刚
已核实刚刚

Staff Software Engineer, Data Enrichment Platform原文

Wayve · Simulation, Evaluation, Validation

官方来源最新
London全职未披露薪资
英文原文

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex…

官方来源职位
首次发现于刚刚
已核实刚刚
India - Bangalore全职未披露薪资
英文原文

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Software Engineering Job Detail…

未出现在监控的职位板上
首次发现于6小时前
已核实6小时前