软件工程师,站点可靠性(SRE)
查看雇主原标题
Software Engineer, Site Reliability (SRE)Sierra · San Francisco, CA · $230k – $390k
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 57/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
- 已披露薪资
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 已披露薪资+15
- 稀有职位+1
- 公司来源健康度+8
该职位未包含:远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
机器翻译关于我们 Sierra 是面向客户的 AI 智能体领域的领先平台,与众多全球知名品牌合作——包括 The GAP、Rocket Mortgage、SoFi、Sutter Health 和 SoftBank——帮助他们转变服务客户和拓展业务的方式。我们主要是一家位于旧金山的线下办公公司,并在北美、欧洲和亚洲设有不断扩展的办公室。 我们以一套价值观为指引,这些价值观是我们行动的核心,也定义了我们的文化:信任、客户至上、匠心、全力以赴,以及在此过程中兼顾家庭的承诺。这些价值观是我们工作的基础,我们致力于在所做的一切中坚守它们。 我们的联合创始人是 Bret Taylor 和 Clay Bavor。Bret 目前担任 OpenAI 的董事会主席。此前,他曾任 Salesforce(该公司收购了他创立的公司 Quip)的联席 CEO 以及 Facebook 的 CTO。Bret 还是 Google 最早的产品经理之一,也是 Google Maps 的联合创造者。在创立 Sierra 之前,Clay 在 Google 工作了 18 年,最近负责领导 Google Labs。更早之前,他发起并领导了 Google 的 AR/VR 项目、Project Starline 和 Google Lens。在那之前,Clay 领导了 Google Workspace 的产品和设计团队。 你将做什么 作为 Sierra 站点可靠性团队的软件工程师,你将负责定义并构建 Sierra AI 驱动基础设施在可靠性、可观测性和可扩展性方面的基础。你将与我们的核心工程和产品团队紧密合作,确保我们的系统具有高可用性、高效率,并为增长而构建。 • 负责 Sierra 的可观测性技术栈——监控、告警、日志和追踪——让工程师能够清晰了解系统健康状况和性能。
• 与产品和平台工程师合作,从第一天起就将系统设计为可靠且可扩展——而不是事后才考虑。
• 使用 Terraform 和现代 DevOps 工具设计并实现可扩展、可靠且安全的云基础设施(AWS)。
• 提升我们 LLM 部署的可靠性和可扩展性,确保稳健、高性能且具有成本效益的运行。
• 主导部署流水线、CI/CD 工具和事件管理流程的改进,以减少停机时间和响应时间。
• 定义 Sierra 的 SRE 实践基础,影响整个工程组织的文化、工具和最佳实践。
你将带来什么 • 5 年以上在复杂 SaaS 或基于云的系统中从事站点可靠性或基础设施工程角色的实操经验。
• 在基础设施层面设计可用性、可扩展性和可靠性的经验
岗位职责
作为 Sierra 站点可靠性团队的软件工程师,你将负责定义并构建 Sierra AI 驱动基础设施在可靠性、可观测性和可扩展性方面的基础。你将与我们的核心工程和产品团队紧密合作,确保我们的系统具有高可用性、高效率,并为增长而构建。 • 负责 Sierra 的可观测性技术栈——监控、告警、日志和追踪——让工程师能够清晰了解系统健康状况和性能。
• 与产品和平台工程师合作,从第一天起就将系统设计为可靠且可扩展——而不是事后才考虑。
• 使用 Terraform 和现代 DevOps 工具设计并实现可扩展、可靠且安全的云基础设施(AWS)。
• 提升我们 LLM 部署的可靠性和可扩展性,确保稳健、高性能且具有成本效益的运行。
• 主导部署流水线、CI/CD 工具和事件管理流程的改进,以减少停机时间和响应时间。
• 定义 Sierra 的 SRE 实践基础,影响整个工程组织的文化、工具和最佳实践。
你将带来什么 • 5 年以上在复杂 SaaS 或基于云的系统中从事站点可靠性或基础设施工程角色的实操经验。
• 在基础设施和应用层面设计可用性、可扩展性和可靠性的经验。
• 在 Terraform、AWS 服务、容器编排和云网络(包括 IAM 和 VPC 架构)方面的深厚经验。
• 在可观测性系统(例如 Prometheus、Grafana、Datadog 或类似系统)方面的扎实背景。
• 与企业客户合作的经验,并熟悉他们的合规和网络需求以及集成模式。
• 能够适应快速变化的环境,并跨产品、ML 和核心工程团队协作。
• 计算机科学或相关领域的学位,或同等专业经验。
更佳条件…… • 具有 LLM 基础设施经验——优化推理性能、管理微调模型或大规模模型部署。
• 有早期初创公司环境的经验,尤其是从零开始定义 SRE 文化和工具。
• 熟悉事件管理自动化或自愈基础设施模式。
我们的价值观 • 信任:我们通过责任感、同理心、质量和响应能力与客户建立信任。我们通过让 AI 更易获取、更安全、更有用来建立对 AI 的信任。我们通过在工作和个人层面彼此支持来建立相互信任,营造一个让我们所有人都能发挥最佳水平的环境。
• 客户至上:我们深入理解客户的业务目标,并始终专注于推动成果,而不仅仅是技术里程碑。公司里的每个人都了解我们的客户并花时间与他们相处。当我们的客户遇到问题时,我们会放下一切去解决它。
• 匠心:我们把细节做对,从页面上的文字到系统架构。我们有良好的品味。当我们发现某些地方不对时,我们会花时间修复它。我们为自己打造的产品感到自豪。我们不断自我反思,以不断自我提升。
• 全力以赴:我们知道我们没有耐心等待的奢侈。我们为胜利而战。我们关心我们的产品是否最好,当它不是时,我们就修复它。当我们失败时,我们会公开讨论且不相互指责,以便下一次成功。
• 家庭:我们知道平衡与全力以赴是可以兼容的,并在我们的行动和流程中体现这一点。我们是最适合为人父母者的科技公司。我们相互支持、相互尊重,并庆祝彼此的个人和职业成就。
福利待遇
我们希望我们的福利能够体现我们的价值观,并为全职员工提供以下福利: • 灵活(无限)带薪休假
• 为你和你的家人提供医疗、牙科和视力福利
• 人寿保险和伤残福利
• 退休计划,取决于就业所在国家/地区
• 育儿假
• 通过 Carrot 提供的生育和家庭建设福利
• 午餐,以及美味零食和咖啡,让你保持精力充沛
• 自由支配福利津贴,让人们能够在最重要的事情上花钱
• 免费阿尔卑斯长号课程
这些福利在 Sierra 的政策中有进一步详细说明,可能因地区而异,并可能随时变更,且须符合任何适用薪酬或福利计划的条款。符合条件的全职员工可以参与 Sierra 的股权计划,但须遵守适用计划和政策的条款。 与我们一起,做你自己 我们正在努力将 AI 的变革力量带给世界上的每一个组织。为此,对我们来说重要的是,我们员工的多样性能够代表我们客户的多样性。我们相信,当我们鼓励、支持和尊重团队中不同的技能和经验时,我们的工作和文化会更好。即使你的经验与职位描述并不完全匹配,我们也鼓励你申请。我们努力以一致的方式评估所有申请人,不考虑种族、肤色、宗教、性别、国籍、年龄、残疾、退伍军人身份、怀孕、性别表达或身份、性取向、公民身份或任何其他受法律保护的类别。
以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。
查看雇主原文
职位描述
About us Sierra is the leading platform for customer-facing AI agents, working with many of the world's biggest brands — including The GAP, Rocket Mortgage, SoFi, Sutter Health, and SoftBank — to transform how they serve customers and grow their businesses. We are primarily an in-person company based in San Francisco, with growing offices across North America, Europe, and Asia. We are guided by a set of values that are at the core of our actions and define our culture: Trust, Customer Obsession, Craftsmanship, Intensity, and a commitment to balancing Family along the way. These values are the foundation of our work, and we are committed to upholding them in everything we do. Our co-founders are Bret Taylor and Clay Bavor . Bret currently serves as Board Chair of OpenAI. Previously, he was co-CEO of Salesforce (which had acquired the company he founded, Quip) and CTO of Facebook. Bret was also one of Google's earliest product managers and co-creator of Google Maps. Before founding Sierra, Clay spent 18 years at Google, where he most recently led Google Labs. Earlier, he started and led Google’s AR/VR effort, Project Starline, and Google Lens. Before that, Clay led the product and design teams for Google Workspace. What you'll do As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI-driven infrastructure. You’ll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth. • Own Sierra’s observability stack—monitoring, alerting, logging, and tracing—to give engineers clear visibility into system health and performance.
• Partner with product and platform engineers to design systems that are reliable and scalable from day one—not as an afterthought.
• Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.
• Improve the reliability and scalability of our LLM deployments, ensuring robust, performant, and cost-effective operation.
• Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.
• Define the foundation of SRE practices at Sierra, influencing culture, tooling, and best practices across the engineering org.
What you'll bring • 5+ years of hands-on experience in Site Reliability or Infrastructure engineering roles for complex SaaS or cloud-based systems.
• Experience designing for availability, scalability, and reliability at both infrastructure
岗位职责
As a Software Engineer on our Site Reliability team at Sierra, you will be responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI-driven infrastructure. You’ll partner closely with our core engineering and product teams to ensure our systems are highly available, efficient, and built for growth. • Own Sierra’s observability stack—monitoring, alerting, logging, and tracing—to give engineers clear visibility into system health and performance.
• Partner with product and platform engineers to design systems that are reliable and scalable from day one—not as an afterthought.
• Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.
• Improve the reliability and scalability of our LLM deployments, ensuring robust, performant, and cost-effective operation.
• Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.
• Define the foundation of SRE practices at Sierra, influencing culture, tooling, and best practices across the engineering org.
What you'll bring • 5+ years of hands-on experience in Site Reliability or Infrastructure engineering roles for complex SaaS or cloud-based systems.
• Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
• Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).
• Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).
• Experience working with enterprise customers and familiarity with their compliance and networking needs along with integration patterns.
• Comfortable working in fast-moving environments and collaborating across product, ML, and core engineering teams.
• Degree in Computer Science or a related field, or equivalent professional experience.
Even better... • Experience with LLM infrastructure — optimizing inference performance, managing fine-tuned models, or large-scale model deployment.
• Past experience in an early-stage startup environment, especially defining SRE culture and tooling from scratch.
• Familiarity with incident management automation or self-healing infrastructure patterns.
Our values • Trust: We build trust with our customers with our accountability, empathy, quality, and responsiveness. We build trust in AI by making it more accessible, safe, and useful. We build trust with each other by showing up for each other professionally and personally, creating an environment that enables all of us to do our best work.
• Customer Obsession: We deeply understand our customers’ business goals and relentlessly focus on driving outcomes, not just technical milestones. Everyone at the company knows and spends time with our customers. When our customer is having an issue, we drop everything and fix it.
• Craftsmanship: We get the details right, from the words on the page to the system architecture. We have good taste. When we notice something isn’t right, we take the time to fix it. We are proud of the products we produce. We continuously self-reflect to continuously self-improve.
• Intensity: We know we don’t have the luxury of patience. We play to win. We care about our product being the best, and when it isn’t, we fix it. When we fail, we talk about it openly and without blame so we succeed the next time.
• Family: We know that balance and intensity are compatible, and we model it in our actions and processes. We are the best technology company for parents. We support and respect each other and celebrate each other’s personal and professional achievements.
福利待遇
We want our benefits to reflect our values and offer the following to full-time employees: • Flexible (unlimited) paid time off
• Medical, dental, and vision benefits for you and your family
• Life insurance and disability benefits
• Retirement plan dependent on country of employment
• Parental leave
• Fertility and family building benefits through Carrot
• Lunch, as well as delicious snacks and coffee to keep you energized
• Discretionary benefit stipend giving people the ability to spend where it matters most
• Free alphorn lessons
These benefits are further detailed in Sierra's policies, may vary by region, and are subject to change at any time, consistent with the terms of any applicable compensation or benefits plans. Eligible full-time employees can participate in Sierra's equity plans subject to the terms of the applicable plans and policies. Be you, with us We're working to bring the transformative power of AI to every organization in the world. To do so, it is important to us that the diversity of our employees represents the diversity of our customers. We believe that our work and culture are better when we encourage, support, and respect different skills and experiences represented within our team. We encourage you to apply even if your experience doesn't precisely match the job description. We strive to evaluate all applicants consistently without regard to race, color, religion, gender, national origin, age, disability, veteran status, pregnancy, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.