工程总监(数据基础设施)
查看雇主原标题
Director of Engineering (Data Infrastructure)Databricks · Bengaluru, 印度
职位信息来自雇主公开的招聘页面。申请前请务必在雇主官网核实详情。
为什么值得关注?
发现指数 52/100,仅依据与该职位一起存储的证据计算。
- 新的雇主官方职位
- 稀有职位匹配
分数构成
- 时效性 (随职位发布时间变化)+18
- 雇主官方来源+15
- 稀有职位+11
- 公司来源健康度+8
该职位未包含:已披露薪资、远程职位、提及签证担保、提及搬迁、未出现在监控的职位板上。
这些理由来自雇主自己的职位描述与我们核实过的来源检查结果。除了已存储的信号之外,我们不做任何推测。
职位描述
机器翻译(P-1490)
Databricks 每天处理数 PB 的数据和数十亿笔交易事件——每一次集群启动、每一次查询执行、每一美元计费,都流经绝不能失败的基础设施。当我们以 99.999% 的准确率要求处理数十亿美元的计费交易,当我们在 100 多个区域每秒摄取数 TB 数据,当五分钟的中断就会造成数百万美元的收入损失和客户信任损失时——基础设施不仅仅重要,它是生死攸关的。我们下一阶段的增长要求灾难恢复系统能够证明可靠性,而不是寄希望于可靠性;要求测试框架能够在部署前发现生产规模的问题;要求正确性保证使计费错误在结构上不可能发生;要求自动化使运营规模随增长呈次线性扩展。
在这个领导岗位上,你将构建使 Databricks 持续增长成为可能的数据基础设施组织。你将在 Bengaluru 建立基础团队,负责保障我们整个变现技术栈计费正确性、运营韧性和零停机恢复的基石系统,同时负责多区域数据摄取、开发者平台和部署自动化,在 PB 级规模上消除摩擦。这不是维护现有系统,而是架构能够使 Databricks 在扩展的同时减轻运营负担的基础设施。你将定义未来十年数据平台的世界级基础设施应该是什么样子。
你将作为我们增长最快的工程中心的创始技术领导者和全球基础设施领导者的战略合作伙伴,迎接这些挑战。除了组建世界级团队外,你还将塑造影响整个公司的架构决策,并倡导基础设施即产品的思维,将基础设施转变为全球范围内的力量倍增器。你将置身于一个诞生于 Apache Spark 和开源文化的工程文化中,在这里技术深度至关重要,基础设施工程师被尊崇为工匠。
理想的候选人曾在这样的公司构建过基础设施组织:在那里,五个九不仅仅是理想,PB 级不是营销话术而是日常,基础设施团队的技术杠杆决定业务是能够扩展还是停滞。你具备辩论数据架构的技术深度、定义多年平台路线图的战略视野、组建顶尖工程师都想加入的团队的领导力,最重要的是,你坚信做对的数据基础设施不仅仅支持业务,它定义了什么是可能的。
你将产生的影响:
• 为每天处理数十亿美元计费交易且对错误零容忍的系统交付基础设施愿景,构建可证明可靠的灾难恢复、能够发现生产环境所见的测试框架、使计费错误在结构上不可能发生的正确性系统,以及能够在故障发生前预测故障的可观测性
• 构建 Bengaluru 的数据基础设施组织,将其打造为印度顶尖基础设施人才的向往之地,招聘多位成为力量倍增器的工程经理,并创造一种文化,让大规模解决困难的分布式系统问题是日常工作
• 负责在 100 多个区域 24/7/365 运行的关键业务系统,在这些系统中即使 99.9% 的正常运行时间也意味着数小时的客户痛苦,推动可靠性改进以防止数百万美元的收入损失,同时通过使系统自愈、自调和自文档化的框架消除运营琐事
• 交付在 Databricks 内部复合工程杠杆的平台:在客户发现之前捕获计费错误的正确性框架、使区域扩展一键完成的部署自动化、无需人工干预即可处理 PB 级数据流的数据集成系统,以及全面覆盖是自动而非英雄主义的测试基础设施
• 将基础设施定位为产品,把内部工程团队视为有 SLA 的客户,衡量采用率和满意度,根据反馈迭代,并证明每一美元的基础设施投资都能在产品速度、可靠性改进或成本降低方面带来倍增回报
你需要具备:
• 14 年以上分布式系统工程经验,其中 6 年以上领导基础设施组织,4 年以上管理经理,且所在公司的基础设施故障意味着直接收入影响、客户升级或监管后果——而你构建了使这些故障变得罕见的系统和团队
• 在 PB 级数据管道和分布式系统可靠性方面的技术深度,能够从“我们应该如何架构多区域灾难恢复”深入到“为什么这个 Kafka 集群表现出这种延迟模式”,同时知道何时该辅导、何时该决策
• 有定义多年基础设施愿景并将其转化为按季度展现价值的顺序交付物的记录,同时朝着架构终态构建,将基础设施投资定位为业务赋能者而非成本中心,并做出随时间复合的构建与购买决策
• 构建 99.999%+ 可靠系统的经验,具备成熟的 SLO/SLI、混沌工程、灾难恢复实践,以及能够在故障发生前预测故障的先进可观测性
• 在高增长环境中扩展基础设施组织的成功经验,曾将工程团队规模翻倍同时保持质量标准,培养工程经理,并创建因问题有趣和文化强大而保持高留存率的团队
• 沟通能力,能够使复杂的基础设施决策对高管清晰易懂(将技术投资转化为业务成果),在没有职权的情况下影响跨职能合作伙伴,在不同时区、不同工作风格的全球团队之间建立信任,并在外部代表 Databricks 的技术品牌
• 计算机科学或工程学士学位;硕士或博士优先。具有 Apache Spark、Delta Lake、大规模数据基础设施、金融科技/计费系统经验,或在超高速增长中领导基础设施的经验者强烈优先
关于 Databricks
Databricks 是数据和 AI 公司。全球超过 20,000 家组织——包括 adidas、AT&T、Bayer、Block、Mastercard、Rivian、Unilever,以及 70% 的财富 500 强——依赖 Databricks Data + AI Platform 构建和扩展数据与 AI 应用、分析和智能体。Databricks 总部位于旧金山,在全球拥有 30 多个办事处,提供统一平台,包括 Genie、Lakebase、Agent Bricks、Lakeflow、Lakehouse 和 Unity Catalog。如需了解更多,请在 LinkedIn、X、YouTube 和 Instagram 上关注 Databricks。
福利待遇
在 Databricks,我们努力提供满足所有员工需求的全面福利和津贴。有关您所在地区所提供福利的具体详情,请点击此处。
我们对多元化和包容性的承诺
在 Databricks,我们致力于培育一个多元化和包容性的文化,让每个人都能脱颖而出。我们非常谨慎地确保我们的招聘实践具有包容性并符合平等就业机会标准。在 Databricks 寻求就业的个人不会因年龄、肤色、残疾、族裔、家庭或婚姻状况、性别认同或表达、语言、国籍、身体和心理能力、政治派别、种族、宗教、性取向、社会经济地位、退伍军人身份及其他受保护特征而受到区别对待。
合规
如果履行工作职责需要访问受出口管制的技术或源代码,雇主可自行决定是否为此类职位申请美国政府许可证,且雇主可能仅基于此原因而拒绝继续处理申请人的申请。
以上内容由机器翻译自动生成,可能存在错误;投递前请以雇主原文为准。
查看雇主原文
职位描述
(P-1490)
Databricks processes petabytes of data and billions of transaction events daily - every cluster launch, every query executed, every dollar billed flows through infrastructure that must never fail. When we process billions in billing transactions with 99.999% accuracy requirements, when we ingest terabytes per second across 100+ regions, when a five-minute outage costs millions in revenue and customer trust - infrastructure isn't just important, it's existential. The next phase of our growth demands disaster recovery systems that prove reliability rather than hope for it, testing frameworks that catch production-scale problems before deployment, correctness guarantees that make billing errors structurally impossible, and automation that scales operations sublinearly with growth.
In this leadership opportunity, you will build the data infrastructure organization that makes Databricks' continued growth possible. You'll establish foundational teams in Bengaluru owning the bedrock systems that guarantee billing correctness, operational resilience, and zero-downtime recovery across our entire monetization stack, alongside multi-region data ingestion, developer platforms, and deployment automation that eliminate friction at petabyte scale. This isn't about maintaining what exists; it's about architecting the infrastructure that enables Databricks to scale while reducing operational burden. You'll define what world-class infrastructure looks like for the next decade of data platforms.
You will pursue these challenges as a founding technical leader in our fastest-growing engineering hub and strategic partner to global infrastructure leaders. In addition to building world-class teams, you will shape architectural decisions that ripple across the company and champion infrastructure-as-product thinking that transforms infrastructure into force multipliers globally. You'll work in an engineering culture born from Apache Spark and open source, where technical depth matters and infrastructure engineers are celebrated as craftspeople.
The perfect candidate has built infrastructure organizations at companies where five nines weren't simply aspirational, where petabyte-scale wasn't marketing but Monday, and where the infrastructure team's technical leverage determined whether the business could scale or stall. You have the technical depth to debate data architecture, the strategic vision to define multi-year platform roadmaps, the leadership craft to build teams that top engineers want to join, and most importantly, the conviction that data infrastructure done right doesn't just support the business; it defines what's possible.
The impact you’ll have:
• Deliver the infrastructure vision for systems processing billions in daily billing transactions with zero tolerance for error, building disaster recovery that's provably reliable, testing frameworks that catch what production sees, correctness systems that make billing errors structurally impossible, and observability that predicts failures before they happen
• Build Bengaluru's data infrastructure organization by establishing it as the destination for India's top infrastructure talent , hiring multiple engineering managers who become force multipliers, and creating a culture where solving hard distributed systems problems at scale is the daily work
• Own business-critical systems operating 24/7/365 across 100+ regions where even 99.9% uptime means hours of customer pain, driving reliability improvements that prevent millions in revenue loss while eliminating operational toil through frameworks that make systems self-healing, self-tuning, and self-documenting
• Ship platforms that compound engineering leverage across Databricks: correctness frameworks that catch billing errors before customers do, deployment automation that makes regional expansion push-button, data integration systems that process petabyte-scale flows without human intervention, and testing infrastructure where comprehensive coverage is automatic, not heroic
• Position infrastructure as product by treating internal engineering teams as customers with SLAs, measuring adoption and satisfaction, iterating based on feedback, and demonstrating that every dollar invested in infrastructure returns multiplicative gains in product velocity, reliability improvements, or cost reductions
What you’ll need:
• 14+ years in distributed systems engineering with 6+ years leading infrastructure organizations and 4+ years managing managers at companies where infrastructure failures meant immediate revenue impact, customer escalations, or regulatory consequences - and you built the systems and teams that made those failures rare
• Technical depth across petabyte-scale data pipelines and distributed systems reliability where you can engage from "how should we architect multi-region disaster recovery" to "why is this Kafka cluster exhibiting this latency pattern" while knowing when to coach versus when to decide
• Track record defining multi-year infrastructure vision and translating it into sequential deliverables that show value quarterly while building toward architectural end states, positioning infrastructure investments as business enablers rather than cost centers, and making build-vs-buy decisions that compound over time
• Experience building 99.999%+ reliable systems with established practices for SLOs/SLIs, chaos engineering, disaster recovery, and sophisticated observability that predicts failures before they happen
• Proven ability to scale infrastructure organizations in high-growth environments where you've doubled engineering while maintaining quality bar, developed engineering managers, and created teams where retention is high because the problems are interesting and the culture is strong
• Communication skills to make complex infrastructure decisions legible to executives (translating technical investments into business outcomes), influence cross-functional partners without authority, build trust across global teams in different timezones with different working styles, and represent Databricks' technical brand externally
• BS in Computer Science or Engineering; MS or Ph.D. preferred. Experience with Apache Spark, Delta Lake, large-scale data infrastructure, fintech/billing systems, or leading infrastructure through hypergrowth strongly preferred
About Databricks
Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn , X , YouTube , and Instagram .
福利待遇
At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here .
Our Commitment to Diversity and Inclusion
At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.
Compliance
If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.