软件工程师,机器学习基础设施
查看雇主原标题
Software Engineer, ML InfrastructureCursor (Anysphere) · San Francisco; New York
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 42/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 稀有职位+1
- 公司来源健康度+8
该职位未包含:已披露薪资、远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
机器翻译我们的使命是实现编码自动化。我们旅程的第一步是打造面向专业程序员的最佳工具,融合创新性研究、设计与工程。我们的组织非常扁平,团队规模小但人才密度高。我们尤其欣赏求真、充满热情且富有创造力的人。我们享受激烈的辩论、疯狂的想法,以及不断交付代码。
岗位职责
ML Infrastructure 团队构建大规模计算、存储和软件基础设施,以支持 Cursor 打造全球最佳智能体编码模型的工作。我们正在寻找有兴趣构建高性能基础设施及配套软件的优秀工程师。该职位将与 ML 研究人员和工程师紧密合作,通过改进我们的训练框架、系统可靠性/性能以及开发者体验来支持他们的工作。 • 与 ML 研究人员合作,提升训练的吞吐量和可靠性
• 与 OEM、云服务提供商及其他方合作,规划并构建前沿的 GPU 基础设施
• 提升计算环境的密度和可扩展性,以支持日益增长的 RL 工作负载
• 创建软件和系统,以自动化构建、监控和运行 GPU 集群
• 构建工作负载调度和数据移动系统,以支持 Cursor 不断增长的训练规模
如果你符合以下条件,可能很适合这个职位 • 在系统与基础设施方向的软件工程方面有深厚背景,尤其是 Python、Typescript、Rust 和 Golang
• 具备分布式存储和网络基础设施经验,尤其是在跨云和裸金属环境的 Linux 系统上
• 接触过大规模系统及其独特挑战,理想情况下涉及数千个节点且资源占用规模可观。
• 在生产环境中使用基础设施即代码和配置管理,覆盖主机和 Kubernetes
加分项 • 具备 Nvidia GPU 搭配 Infiniband 或 RoCE 的运维经验,尤其是 Blackwell 和 Hopper 级硬件
• 接触过 Ray、Slurm 或其他常见的计算与运行时调度器
#LI-DNI
以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。
查看雇主原文
职位描述
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
岗位职责
The ML Infrastructure team builds large-scale compute, storage, and software infrastructure to support Cursor’s work building the world’s best agentic coding model. We’re looking for strong engineers who are interested in building high-performance infrastructure and the software to support it. This role works closely with ML researchers and engineers to enable their work through improvements to our training framework, systems reliability/performance, and developer experience. • Collaborate with ML researchers to improve the throughput and reliability of training
• Work with OEMs, cloud service providers, and others to plan and build cutting-edge GPU infrastructure
• Improve the density and scalability of compute environments to enable increasingly large RL workloads
• Create software and systems to automate building, monitoring, and running GPU clusters
• Build workload scheduling and data movement systems to support Cursor’s growing training footprint
You may be a fit if • A strong background in systems and infrastructure-focused software engineering, particularly in Python, Typescript, Rust, and Golang
• Experience with distributed storage and networking infrastructure, particularly on Linux systems across cloud and bare metal environments
• Exposure to large-scale systems and their unique challenges, ideally across thousands of nodes with significant resource footprints.
• Production use of infrastructure-as-code and configuration management, across hosts and Kubernetes
Nice to have • Operational exposure to Nvidia GPUs with Infiniband or RoCE, particularly with Blackwell and Hopper-class hardware
• Exposure to Ray, Slurm, or other common compute and runtime schedulers
#LI-DNI