跳到主要内容
OOfficialJobs
菜单
官方来源官方来源职位

资深 / 高级机器学习工程师,强化学习

机器翻译
查看雇主原标题Staff / Senior Machine Learning Engineer, Reinforcement Learning

Wayve · Sunnyvale

职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。

为什么值得关注?

发现指数 42/100,仅依据与该职位一起存储的证据计算。

42/100 发现指数
  • 新的雇主官方职位

分数构成

  • 时效性 (随职位发布时间变化)+18
  • 雇主官方来源+15
  • 稀有职位+1
  • 公司来源健康度+8

该职位未包含:已披露薪资、远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。

这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。

职位描述

英文原文

该职位由雇主以英文发布,暂无中文版本,下面完整显示英文原文。 查看官方职位页面.

职位描述

About us

Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.

Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.

In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.

At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.

Make Wayve the experience that defines your career!

The role

As a Senior / Staff Machine Learning Engineer in Wayve's AV Core organisation, you will advance reinforcement learning methods for end-to-end driving models. You will identify where learning from reward or feedback can improve beyond behavior cloning, then take promising ideas from design through large-scale experiments, rigorous evaluation, and integration into our best driving models.

Driving Core team develops the learning methods that turn diverse driving data into robust closed-loop behavior. You will be a technical owner for reinforcement learning within the group, working closely with researchers and engineers across AV Core, Simulation, Evaluation, and Product Engineering. Success means producing measurable improvements in driving behavior.

Core Model Safety team develops the core model competencies that enable safe, driverless operation. You will lead the technical direction and delivery of a learned emergency trajectory model for low-frequency, high-consequence maneuvers such as evasive steering and emergency braking. You will take the programme from problem definition through modelling, evaluation, integration, and evidence for deployment.

Key responsibilities

• Shape and execute the reinforcement learning roadmap for Driving Core / Core Model Safety, selecting problems and methods against clear behavioral gaps and measurable success criteria.

• Develop and evaluate post-behavior-cloning optimization methods, including offline and off-policy reinforcement learning as well as other reward-guided approaches; design the regularization, data strategy, and diagnostics needed to make policies reliably better.

• Help improve the reward models and related learning signals used to train and evaluate driving policies, working with partner teams to strengthen their quality, scalability, and downstream usefulness.

• Build robust training and experimentation workflows using large-scale driving data; diagnose distribution shift, objectiv

岗位职责

As a Senior / Staff Machine Learning Engineer in Wayve's AV Core organisation, you will advance reinforcement learning methods for end-to-end driving models. You will identify where learning from reward or feedback can improve beyond behavior cloning, then take promising ideas from design through large-scale experiments, rigorous evaluation, and integration into our best driving models.

Driving Core team develops the learning methods that turn diverse driving data into robust closed-loop behavior. You will be a technical owner for reinforcement learning within the group, working closely with researchers and engineers across AV Core, Simulation, Evaluation, and Product Engineering. Success means producing measurable improvements in driving behavior.

Core Model Safety team develops the core model competencies that enable safe, driverless operation. You will lead the technical direction and delivery of a learned emergency trajectory model for low-frequency, high-consequence maneuvers such as evasive steering and emergency braking. You will take the programme from problem definition through modelling, evaluation, integration, and evidence for deployment.

• Shape and execute the reinforcement learning roadmap for Driving Core / Core Model Safety, selecting problems and methods against clear behavioral gaps and measurable success criteria.

• Develop and evaluate post-behavior-cloning optimization methods, including offline and off-policy reinforcement learning as well as other reward-guided approaches; design the regularization, data strategy, and diagnostics needed to make policies reliably better.

• Help improve the reward models and related learning signals used to train and evaluate driving policies, working with partner teams to strengthen their quality, scalability, and downstream usefulness.

• Build robust training and experimentation workflows using large-scale driving data; diagnose distribution shift, objective misspecification, optimization instability, and data or evaluation bias.

• Define evidence across offline metrics, open-loop tests, closed-loop simulation, and on-road evaluation, and distinguish genuine policy improvement from benchmark overfitting.

• Productionize successful methods in the shared ML stack, communicate decisions and results clearly, and raise the technical bar through design reviews, code reviews, and mentoring.

任职要求

In order to set you up for success as a Staff / Senior Machine Learning Engineer at Wayve, we’re looking for the following skills and experience.

• A strong track record developing and experimentally validating reinforcement learning or closely related sequential decision-making methods on complex, high-dimensional problems.

• Deep understanding of modern reinforcement learning fundamentals, including policy and value learning, off-policy learning, function approximation, distribution shift, and the failure modes of learned objectives.

• Hands-on experience with behaviour cloning, reinforcement learning, or related methods.

• Proficiency in Python and PyTorch, with strong software engineering practices and hands-on experience building reliable machine learning training and evaluation systems.

• Excellent experimental judgement: able to turn an ambiguous behavioral problem into falsifiable hypotheses, useful metrics, disciplined ablations, and clear technical decisions.

• Senior-level ownership and collaboration: able to lead a substantial technical area, work across research and engineering boundaries, and bring others along through clear written and verbal communication.

Desirable

• Experience with offline reinforcement learning, imitation learning, reward modeling, preference learning, or post-training of large neural policies.

• Experience in autonomous vehicles, robotics, control, or another domain where policies interact with safety-critical physical systems, including an understanding of motion planning, vehicle dynamics, control, or collision avoidance.

• Experience with closed-loop simulation, off-policy evaluation, uncertainty or calibration, and evaluation under rare or shifted conditions.

• Experience training multimodal, transformer-based, or generative policy models at scale.

• Proficiency in C++, CUDA, distributed training, or performance optimization for production machine learning systems.

This is a full-time role based in our office in Sunnyvale. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. The reasonably estimated salary for this role ranges from $ 311,850 to $ 389,400, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience. Wayve is committed to creating an inclusive interview experience. If you require any accommodations or adjustments to participate fully in our interview process, please let us know.

We understand that everyone has a unique set of skills and experiences and that not everyone will meet all of the requirements listed above. If you’re passionate about self-driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply.

At Wayve we're committed to creating a diverse, fair and respectful culture that is inclusive of everyone based on their unique skills and perspectives, and regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, veteran status, pregnancy or related condition (including breastfeeding) or any other basis as protected by applicable law.

For more information visit Careers at Wayve.

To learn more about what drives us, visit Values at Wayve

For US candidates only, please visit E-Verify Notice and Participation and Right to Work

DISCLAIMER: We will not ask about marriage or pregnancy, care responsibilities or disabilities in any of our job adverts or interviews. However, we do look to capture information about care responsibilities, and disabilities among other diversity information as part of an optional DEI Monitoring form to help us identify areas of improvement in our hiring process and ensure that the process is inclusive and non-discriminatory.

Wayve 的更多职位

公司主页

Trainer原文

Wayve · Fleet Operations

官方来源最新
德国全职未披露薪资
英文原文

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex…

官方来源职位
首次发现于1小时前
已核实1小时前

Technical Configuration Manager原文

Wayve · Product & Delivery

官方来源最新
Germany; London全职未披露薪资
英文原文

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex…

官方来源职位
首次发现于1小时前
已核实1小时前
官方来源最新
Sunnyvale全职$209k – $260k
英文原文

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex…

官方来源职位
首次发现于1小时前
已核实1小时前
官方来源最新
Detroit全职$85k – $750k
英文原文

About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex…

官方来源职位
首次发现于1小时前
已核实1小时前

其他公司的相似职位

搜索这类职位
Mexico - Mexico City全职未披露薪资

为了获得最佳的候选人体验,请考虑在12个月内最多申请3个职位,以确保您不会重复投递。 职位类别 市场营销与传播 职位详情 关于Salesforce Salesforce是排名第一的AI CRM,在这里,人类与智能体携手推动客户成功。在这里,雄心与行动相遇。科技与信任相遇。创新不是流行语——而是一种生活方式。我们所熟知的工作世界正在改变,我们正在寻找Trailblazer,他们热衷于通过AI改善商…

官方来源职位
首次发现于2小时前
已核实2小时前
官方来源最新
San Francisco远程全职$221k – $278k

关于团队 AI Architect 团队与各组织合作,将 OpenAI 最强大的模型转化为有意义的现实世界影响力。我们与医疗健康与生命科学领域的组织合作,识别 AI 可以创造价值的环节,设计安全且可扩展的解决方案,并帮助这些解决方案从早期探索走向持续的生产环境采用。团队汇聚了技术战略、客户合作和实际部署方面的专业能力,并与销售、产品、工程、研究和专业交付团队紧密协作。 关于该职位 作为 AI A…

官方来源职位
首次发现于2小时前
已核实2小时前
官方来源最新
San Francisco全职未披露薪资

我们是谁 关于 Stripe Stripe 是一个面向企业的金融基础设施平台。数百万家公司——从全球最大的企业到最有抱负的初创公司——都在使用 Stripe 来接受付款、增长收入并加速新的商业机会。我们的使命是提升互联网的 GDP,而我们面前还有大量工作要做。这意味着你拥有一个前所未有的机会,在从事职业生涯中最重要工作的同时,让全球经济触手可及。 关于团队 实验项目团队会快速测试 Stripe…

官方来源职位
首次发现于2小时前
已核实2小时前
New York, New York, 美国全职$164k – $215k

你好,我们是Oscar。我们正在招聘一名高级数据科学家,AI赋能方向,加入我们的数据团队。 Oscar是第一家围绕全栈技术平台构建、并始终不渝地专注于服务会员的健康保险公司。我们在2012年创立Oscar,旨在打造一家我们自己想要的健康保险公司——一家像家庭医生一样行事的公司。

官方来源职位
首次发现于2小时前
已核实2小时前