软件工程师,RL Data
查看雇主原标题
Software Engineer, RL DataCursor (Anysphere) · San Francisco
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 42/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 稀有职位+1
- 公司来源健康度+8
该职位未包含:已披露薪资、远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
英文原文该职位由雇主以英文发布,暂无中文版本,下面完整显示英文原文。 查看官方职位页面.
职位描述
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code. Software Engineer, Reinforcement learning • SpaceXAI is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.
岗位职责
• As a Software Engineer on the RL Data team at SpaceXAI, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.
• Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.
• Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.
• Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
• Partnering with research on whether a dataset is actually teaching the thing we think it is.
You may be a fit if • You write careful, fast code and have strong software engineering fundamentals.
• You like setting tasks: breaking a fuzzy capability into something concrete you can measure.
• You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.
• You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.