机器学习工程经理
查看雇主原标题
Engineering Manager, MLCursor (Anysphere) · San Francisco; New York
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 44/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 稀有职位+3
- 公司来源健康度+8
该职位未包含:已披露薪资、远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
机器翻译我们的使命是实现编码自动化。我们旅程的第一步是为专业程序员打造最好的工具,结合富有创造力的研究、设计和工程。我们的组织非常扁平,团队规模小且人才密集。我们尤其喜欢求真、充满热情且富有创造力的人。我们享受激烈的辩论、疯狂的想法以及交付代码。
岗位职责
你将领导一个工程师团队,构建用于训练、测试和评估我们模型的基础设施。这是 SpaceXAI 中少数几个基础设施与模型行为直接交汇的岗位之一:当出现问题时,很少能一眼看出是系统缺陷,还是模型完全按照训练目标在行事,而你的团队必须擅长在修复之前分辨出两者的区别。 你将为我们如何大规模训练和评估模型设定技术方向,保持与代码足够接近以便与团队一起调试,并每天与研究人员合作,把延迟、质量和成本方面的权衡转化为真正会被构建出来的基础设施。我们正在为这个岗位招聘不同职责范围的人选,具体取决于经验和准备承担的问题规模。 示例项目包括.. • 构建 rollout 基础设施,让研究人员能够大规模运行 RL 实验,而不必与底层管道较劲。
• 设计评估流水线,在回归问题发布前将其捕获,并为研究人员提供快速、可信的信号,判断某项改动是否真正带来了帮助。
• 负责模型训练和测试所处的环境:沙箱化、可复现,并且足够快,使迭代速度不会成为瓶颈。
• 为团队衡量质量和进展的方式带来严谨性,尤其是在“是否发布”并不等同于“是否有效”的地方。
• 与研究团队合作,将模型层面的权衡(延迟、质量、成本)转化为具体的基础设施决策。
• 招聘并发展团队:寻找、面试并招揽优秀的基础设施工程师,同时通过辅导、指导和具有高杠杆效应的项目分配来培养你的工程师。
如果你符合以下条件,可能很适合 • 你曾领导工程团队构建用于在生产环境中训练、评估或服务 ML 模型的基础设施。
• 你具备扎实的基础设施和分布式系统基础:你了解真实负载下的可靠性和性能是什么样,而不仅仅是设计文档中的样子。
• 你真心希望保持技术能力:你乐于编写代码、深入审查 PR,并使用像 Cursor 本身这样的工具来快速推进。
• 你能够在模糊环境中自如应对:你会提出正确的问题,在信息不完整的情况下做出合理决策,并帮助团队找到前进的道路。
• 你有招聘和培养工程师的记录,且他们在这个阶段比你当年更优秀。
• 你能流利地与研究人员讨论模型行为,与工程师讨论系统设计,并且你知道什么时候问题其实属于另一个团队。
• 加分项:具备 RL 训练基础设施、评估框架,或为模型训练或测试构建和维护模拟环境的实操经验。
以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。
查看雇主原文
职位描述
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
岗位职责
You will lead a team of engineers building the infrastructure used to train, test, and evaluate our models. This is one of the few places at SpaceXAI where infrastructure and model behavior meet directly: when something breaks, it's rarely obvious whether it's a systems bug or the model doing exactly what it was trained to do, and your team has to be good at telling the difference before they can fix it. You'll set technical direction for how we train and evaluate models at scale, stay close enough to the code to debug alongside your team, and work daily with researchers to turn tradeoffs in latency, quality, and cost into infrastructure that actually gets built. We're hiring across a range of scope for this role, depending on experience and the size of problem you're ready to own. Example projects include.. • Building the rollout infrastructure that lets researchers run RL experiments at scale without fighting the plumbing.
• Designing eval pipelines that catch regressions before they ship, and give researchers fast, trustworthy signal on whether a change actually helped.
• Owning the environments in which models are trained and tested: sandboxed, reproducible, and fast enough that iteration speed isn't the bottleneck.
• Bringing rigor to how the team measures quality and progress, in places where "did it ship" isn't the same as "did it work?"
• Partnering with research to translate model-level tradeoffs (latency, quality, cost) into concrete infrastructure decisions.
• Hiring and growing the team: sourcing, interviewing, and closing exceptional infrastructure engineers, while developing your engineers through coaching, mentorship, and high-leverage project assignments.
You may be a fit if • You've led engineering teams building infrastructure that trains, evaluates, or serves ML models in production.
• You have strong infrastructure and distributed systems fundamentals: you know what reliability and performance look like under real load, not just in a design doc.
• You genuinely want to stay technical: you're comfortable writing code, reviewing PRs with depth, and using tools like Cursor itself to move fast.
• You’re comfortable operating in ambiguity: you ask the right questions, make sound decisions with incomplete information, and help the team find a path forward.
• You have a track record of hiring and developing engineers who are better than you were at their stage.
• You can talk fluently with researchers about model behavior and with engineers about systems design, and you know when a problem is actually the other team's.
• Bonus: hands-on experience with RL training infrastructure, eval frameworks, or building and maintaining simulated environments for model training or testing.