用户研究员,AI 评估
查看雇主原标题
User Researcher, AI EvaluationsNotion · San Francisco, California; New York, New York
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 60/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
- 远程职位
- 稀有职位匹配
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 远程职位+8
- 稀有职位+11
- 公司来源健康度+8
该职位未包含:已披露薪资、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
机器翻译我们是谁 Notion 是协作式 AI 工作空间,团队和智能体在这里共同思考。我们正在打造一个统一的地方,让你的知识、项目、会议和 AI 工具并排共存,从而让工作更快速、更清晰、更少割裂。数百万个人、小型团队和大型公司都在 Notion 上开展工作。 Notinos(我们的员工)是实现这一未来工作方式的零号客户。我们重视匠心、打造经久耐用的东西,并坚信伟大的工作从根本上仍然是人的工作。我们的目标不是发布下一个功能。每一支 Notinos 团队都在努力为 AI 时代人类如何协作树立标准。从构建企业的记录系统,到创建和管理 AI 智能体,再到自动化处理繁琐工作,我们深切关注如何让客户有更多时间投入到他们毕生的事业中。 关于该职位: 我们正在寻找一位经验丰富的用户体验研究员,来定义并扩展我们评估 Notion AI 驱动体验的方式——不仅关注模型输出质量的“好”是什么样,也关注端到端产品体验中的“好”是什么样,即人们如何发现、设定目标、委派工作、审阅结果,并随着时间推移与 AI 建立信任。
该职位处于研究技艺与评估运营的交汇点:你将开展研究,揭示用户的心智模型、期望以及失败/恢复行为,然后将这些洞察转化为可复用的评分标准、工作流程和衡量方法,供产品、设计、工程和数据科学团队一致应用。
该职位可设在旧金山或纽约市。我们在周一、周二和周四到办公室工作(我们的锚定日),因为面对面时我们最能一起思考和构建。我们正在寻找一位期待在这些日子与团队并肩工作的人。 你将实现的目标: • 定义“好”是什么样(框架与评分标准):建立清晰、可复用的评估标准,反映真实用户期望——有用性、信任、语气、控制感和透明度。你将把定性洞察转化为评分指导,使其能够跨团队并随时间推移一致应用。
• 开展周期性评估(纵向与特定功能):开展周期性的纵向和特定功能调研与研究,依据既定评分标准衡量体验质量随时间的变化。主导定性研究、并排对比以及人在回路中的评估工作,以加深对体验在何处失效以及如何改进的理解。你将帮助团队发现回退、对标改进,并理解期望何时发生变化。
• 将评估锚定在真实工作流程中(上下文 > 孤立反馈):确保评估反映待完成的任务、用户意图以及完整交互旅程(目标设定、委派、审阅、迭代),而不仅仅是脱离上下文的点赞/点踩。你将帮助团队理解谁在评估、他们想做什么,以及输出为何成功或失败。
• 识别失败模式与恢复行为(护栏):揭示整个系统中的故障、回退和边缘情况——从模型行为到 UI 和集成——并研究人们如何发现问题、纠正问题并继续工作。你将把这些洞察转化为关于护栏、修复和优先级排序的可执行指导。
• 与合作伙伴一起将评估运营化(流程与工具):与产品、设计、工程和数据科学紧密协作,对齐目标用例,并构建可扩展的评估循环(人在回路中的审阅、纵向研究,以及将自动化/LLM 评审方法与人类判断进行校准)。
岗位职责
我们正在寻找一位经验丰富的用户体验研究员,来定义并扩展我们评估 Notion AI 驱动体验的方式——不仅关注模型输出质量的“好”是什么样,也关注端到端产品体验中的“好”是什么样,即人们如何发现、设定目标、委派工作、审阅结果,并随着时间推移与 AI 建立信任。
该职位处于研究技艺与评估运营的交汇点:你将开展研究,揭示用户的心智模型、期望以及失败/恢复行为,然后将这些洞察转化为可复用的评分标准、工作流程和衡量方法,供产品、设计、工程和数据科学团队一致应用。
该职位可设在旧金山或纽约市。我们在周一、周二和周四到办公室工作(我们的锚定日),因为面对面时我们最能一起思考和构建。我们正在寻找一位期待在这些日子与团队并肩工作的人。 你将实现的目标: • 定义“好”是什么样(框架与评分标准):建立清晰、可复用的评估标准,反映真实用户期望——有用性、信任、语气、控制感和透明度。你将把定性洞察转化为评分指导,使其能够跨团队并随时间推移一致应用。
• 开展周期性评估(纵向与特定功能):开展周期性的纵向和特定功能调研与研究,依据既定评分标准衡量体验质量随时间的变化。主导定性研究、并排对比以及人在回路中的评估工作,以加深对体验在何处失效以及如何改进的理解。你将帮助团队发现回退、对标改进,并理解期望何时发生变化。
• 将评估锚定在真实工作流程中(上下文 > 孤立反馈):确保评估反映待完成的任务、用户意图以及完整交互旅程(目标设定、委派、审阅、迭代),而不仅仅是脱离上下文的点赞/点踩。你将帮助团队理解谁在评估、他们想做什么,以及输出为何成功或失败。
• 识别失败模式与恢复行为(护栏):揭示整个系统中的故障、回退和边缘情况——从模型行为到 UI 和集成——并研究人们如何发现问题、纠正问题并继续工作。你将把这些洞察转化为关于护栏、修复和优先级排序的可执行指导。
• 与合作伙伴一起将评估运营化(流程与工具):与产品、设计、工程和数据科学紧密协作,对齐目标用例,并构建可扩展的评估循环(人在回路中的审阅、纵向研究,以及将自动化/LLM 评审方法与人类判断进行校准)。
任职要求
• 能够将洞察转化为衡量方式:你擅长将“软性”用户期望(信任、语气、有用性、清晰度)转化为具体的评分标准、评分指南和可观察的指标。
• AI 熟练度与系统思维:你对 AI 产品充满好奇并亲自动手实践,能够推理模型行为、不确定性和系统约束如何塑造用户体验。你也有评估 AI 赋能产品(LLM、智能体、生成式 UI/工作流自动化)的经验,并与数据科学/机器学习合作伙伴在衡量策略和评估工具方面开展协作。
• 清晰的沟通与结果导向:你能够让多元合作伙伴围绕共同的质量定义达成一致,并创建使团队能够一致行动的工作成果。你能针对不同受众调整叙事方式,将研究与业务成果联系起来,并推动后续落实,使洞察转化为产品变革。
• 扎实的用户体验研究技艺(定量 + 定性):你能为问题选择合适的方法——访谈、对标、调研、实验——并综合成可执行的指导。你也能果断排定优先级,在模糊中推进,并在需要时平衡快速迭代与深入探究。
• 在快速变化环境中的务实精神:你能果断排定优先级,在模糊中推进,并在需要时平衡快速迭代与深入探究。
加分项: • 熟悉 LLM 作为评审的方法、面向评估者的提示设计,或“黄金数据集”创建
• 有使用 AI 研究工具进行快速综合与沟通的经验(例如 Dovetail、Listen Labs、Maze、Outset 等),以及使用 Braintrust 等 AI 可观测性工具的经验
• 有使用数据查询语言(例如 SQL)、脚本语言(例如 Python)或统计/数学软件(例如 R、SAS、Matlab 等)的经验
• 人机交互、心理学、行为科学、人类学、社会学或相关领域的硕士或博士学位
• 你熟悉 Douglas Engelbart、Alan Kay、Bret Victor 等计算领域先驱的工作——并理解我们为何是他们的忠实粉丝。
Notion 致力于提供极具竞争力的现金薪酬、股权和福利。该职位提供的薪酬将基于多种因素,例如地点、职位范围与复杂度,以及候选人的经验与专长,并可能不同于下方提供的范围。对于设在旧金山或纽约市的职位,该职位估计的基本薪资范围为每年 $196,000-$230,000。 点击“提交申请”,即表示我理解并同意 Notion 及其关联公司和子公司将根据 Notion 的全球招聘隐私政策和 NYLL 144 收集和处理我的信息。 #LI-Onsite
关于 AI 的说明 并非每个职位都需要深厚的 AI 专业知识,但我们期望每一位 Notino 都具备求知欲,热衷于动手尝试和探索,并期待将 AI 作为工作中的真正协作者。对于某些职位,AI 熟练度是核心要求——如果是这种情况,我们会在任职资格中明确说明。在这里如鱼得水的人不会把 AI 当作新奇事物。他们用它来更好地思考,并让他人更容易在他们的工作基础上继续构建。 平等机会与便利安排 我们聘用来自广泛背景的优秀人才。如果你对该职位感到兴奋,但并非满足每一条要求,我们仍然鼓励你申请。Notion 是提供平等机会的雇主,不会基于任何受法律保护的特征进行歧视。根据适用法律,我们会考虑有逮捕和定罪记录的合格申请人。Notion 在申请过程中提供合理的便利安排;如果你需要,请告知你的招聘人员。 Notion 自豪地成为提供平等机会的雇主。我们在招聘或任何雇佣决定中不会基于种族、肤色、宗教、国籍、年龄、性别(包括怀孕、分娩或相关医疗状况)、婚姻状况、血统、身体或精神残疾、遗传信息、退伍军人身份、性别认同或表达、性取向或其他适用的受法律保护特征进行歧视。Notion 会根据适用的联邦、州和地方法律考虑有犯罪历史的合格申请人。Notion 也致力于在我们的求职申请流程中为符合条件的残障人士和残障退伍军人提供合理的便利安排。如果你因残障需要协助或便利安排,请告知你的招聘人员。
以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。
查看雇主原文
职位描述
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for the end-to-end product experience where people discover, set goals, delegate work, review results, and build trust over time with AI.
This role sits at the intersection of research craft and evaluation operations: you’ll run studies that uncover user mental models, expectations, and failure/recovery behaviors, then translate those insights into reusable rubrics, workflows, and measurement approaches that product, design, engineering, and data science can apply consistently.
This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: • Define what “good” looks like (frameworks & rubrics): Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness, trust, tone, control, and transparency. You’ll translate qualitative insight into scoring guidance that can be applied consistently across teams and over time.
• Run recurring evals (longitudinal & feature-specific): Run recurring longitudinal and feature-specific surveys and studies to measure experience quality over time against defined rubrics. Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluation efforts to deepen understanding of where experiences break down and how they can improve. You’ll help teams spot regressions, benchmark improvements, and understand when expectations shift.
• Anchor evaluation in real workflows (context > isolated feedback): Ensure evals reflect jobs-to-be-done, user intent, and the full interaction journey (goal setting, delegation, review, iteration), not just decontextualized thumbs up/down. You’ll help teams understand who is evaluating, what they’re trying to do, and why outputs succeed or fail.
• Identify failure modes & recovery behavior (guardrails): Uncover breakdowns, regressions, and edge cases across the system—from model behavior to UI and integrations—and study how people notice issues, correct them, and continue their work. You’ll turn these insights into actionable guidance for guardrails, fixes, and prioritization.
• Operati
岗位职责
We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for the end-to-end product experience where people discover, set goals, delegate work, review results, and build trust over time with AI.
This role sits at the intersection of research craft and evaluation operations: you’ll run studies that uncover user mental models, expectations, and failure/recovery behaviors, then translate those insights into reusable rubrics, workflows, and measurement approaches that product, design, engineering, and data science can apply consistently.
This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: • Define what “good” looks like (frameworks & rubrics): Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness, trust, tone, control, and transparency. You’ll translate qualitative insight into scoring guidance that can be applied consistently across teams and over time.
• Run recurring evals (longitudinal & feature-specific): Run recurring longitudinal and feature-specific surveys and studies to measure experience quality over time against defined rubrics. Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluation efforts to deepen understanding of where experiences break down and how they can improve. You’ll help teams spot regressions, benchmark improvements, and understand when expectations shift.
• Anchor evaluation in real workflows (context > isolated feedback): Ensure evals reflect jobs-to-be-done, user intent, and the full interaction journey (goal setting, delegation, review, iteration), not just decontextualized thumbs up/down. You’ll help teams understand who is evaluating, what they’re trying to do, and why outputs succeed or fail.
• Identify failure modes & recovery behavior (guardrails): Uncover breakdowns, regressions, and edge cases across the system—from model behavior to UI and integrations—and study how people notice issues, correct them, and continue their work. You’ll turn these insights into actionable guidance for guardrails, fixes, and prioritization.
• Operationalize evaluation with partners (process & tooling): Collaborate closely with Product, Design, Engineering, and Data Science to align on target use cases and build scalable evaluation loops (human-in-the-loop review, longitudinal studies, and calibration of automated/LLM-judge approaches against human judgment).
任职要求
• Ability to operationalize insight into measurement: You’re comfortable turning “soft” user expectations (trust, tone, usefulness, clarity) into concrete rubrics, scoring guidelines, and observable metrics.
• AI fluency and systems thinking: You’re curious and hands-on with AI products, and can reason about how model behavior, uncertainty, and system constraints shape user experience. You also have experience evaluating AI-enabled products (LLMs, agents, generative UI/workflow automation) and working with Data Science/ML partners on measurement strategy and evaluation tooling.
• Clear communication and impact orientation: You can align diverse partners around shared definitions of quality and create artifacts that enable teams to act consistently. You tailor storytelling to different audiences, connect research to business outcomes, and drive follow-through so insights translate into product change.
• Strong UX research craft (quant + qual): You can choose the right methods for the question— interviews, benchmarking, surveys, experiments—and synthesize into actionable guidance. You also can prioritize ruthlessly, work through ambiguity, and balance scrappy iteration with deep dives when needed.
• Pragmatism in fast-moving environments: You can prioritize ruthlessly, work through ambiguity, and balance scrappy iteration with deep dives when needed.
Nice to Haves: • Familiarity with LLM-as-judge methods, prompt design for evaluators, or “golden dataset” creation
• Experience using AI research tooling for rapid synthesis and communication (e.g., Dovetail, Listen Labs, Maze, Outset, etc.), as well as AI observability tooling like Braintrust
• Experience using data querying languages (e.g., SQL), scripting languages (e.g., Python), or statistical/mathematical software (e.g., R, SAS, Matlab, etc.)
• Master’s or PhD in HCI, Psychology, Behavioral Science, Anthropology, Sociology, or a related field
• You’re familiar with the work of computing heroes like Douglas Engelbart, Alan Kay, Bret Victor, etc. — and understand why we're big fans.
Notion is committed to providing highly competitive cash compensation, equity, and benefits. The compensation offered for this role will be based on multiple factors such as location, the role’s scope and complexity, and the candidate’s experience and expertise, and may vary from the range provided below. For roles based in San Francisco or New York City, the estimated base salary range for this role is $196,000-$230,000 per year. By clicking “Submit Application”, I understand and agree that Notion and its affiliates and subsidiaries will collect and process my information in accordance with Notion’s Global Recruiting Privacy Policy and NYLL 144 . #LI-Onsite
A Note on AI You don’t need deep AI expertise for every role, but we do expect every Notino to be intellectually curious, drawn to tinkering and discovery, and excited to use AI as a real collaborator in their work. For some roles, AI fluency is a core requirement — when that’s the case, we'll say so explicitly in the qualifications. People who thrive here don’t treat AI as a novelty. They use it to think better, and make their work easier for others to build on. Equal Opportunity & Accommodations We hire talented people from a wide range of backgrounds. If you’re excited about this role but don’t meet every bullet, we still encourage you to apply. Notion is an equal opportunity employer and does not discriminate on the basis of any legally protected characteristic. Consistent with applicable law, we will consider for employment qualified applicants with arrest and conviction records. Notion provides reasonable accommodations during the application process; if you need one, please let your recruiter know. Notion is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristic. Notion considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Notion is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please let your recruiter know.