Senior Software Engineering Manager FinOps Platform Services
Pleasanton, CA - USA
Job Summary
What you get to do in this role:
Platform Ownership & Operations
Own the operational health and reliability of Trino Lightdash Coder Jupyter Redash Hive Metastore and Nessie across development and production environments.
Establish and maintain SLOs for platform availability query performance and workspace provisioning. Build the dashboards and alerting to track them.
Own Trino cluster operations end to end including deployment scaling upgrades performance tuning resource group management query optimization support and user access controls.
Drive the platform upgrade and patching cadence balancing stability with staying current on security fixes and feature releases across all services.
Build runbooks on-call processes and incident-response practices so the team can respond to and resolve production issues quickly and learn from them.
Ensure platform security across all services including access controls authentication (SSO/OIDC integration) secrets management and audit logging.
Platform Evolution & Roadmap
Lead the migration from Hive Metastore to Nessie as the versioned Iceberg catalog delivering Git-like branching semantics safe multi-writer coordination and auditable catalog history.
Drive Lightdash platform improvements including version upgrades performance optimization row-level security configuration and the governed self-service analytics experience.
Evolve the Coder platform through workspace template lifecycle management resource policies idle-stop tuning and onboarding new users and use cases including AI coding agents.
Own the Jupyter and Redash platforms ensuring availability scaling integration with Trino and the lakehouse and user lifecycle management.
Evaluate and adopt new open-source technologies where they raise the platforms ceiling or reduce operational burden.
People Leadership
Manage mentor and grow a team of 3 to 5 platform engineers. Set clear expectations provide regular feedback and create career development paths.
Hire and build the team to match the platforms growing scope and user base.
Foster a culture of operational excellence automation over toil and blameless incident retrospectives.
Set engineering standards for how the team builds deploys monitors and documents platform services.
Collaboration & Stakeholder Management
Partner with the DevOps/infrastructure team on Kubernetes capacity networking storage and CI/CD pipeline needs for your platform services.
Serve as the platform liaison to data engineers analysts and FinOps practitioners. Understand their workflows gather feedback and prioritize improvements that unblock them.
Collaborate with the Data Platform and Data Governance teams to ensure platform services align with enterprise standards for security lineage and access control.
Support the broader Cloudera-to-lakehouse migration by ensuring Trino Nessie and the catalog layer are production-ready for migrated workloads.
Apply AI/ML tooling where it accelerates platform operations monitoring or user support.
What success looks like
Platform services meet their SLOs consistently and the team has the observability and processes to detect and resolve issues before users are affected.
Trino queries perform reliably at scale with well-managed resource groups and a clear upgrade cadence.
The Hive Metastore to Nessie migration is planned sequenced and executing without disruption to downstream users.
Lightdash and Coder are stable current and adopted broadly across the organization with minimal friction for new users.
The team is healthy growing and operating with clear ownership automation and documentation.
Internal users trust the platform and rarely lose productive time to platform instability.
Qualifications :
To be successful in this role you have:
Experience leveraging or critically thinking about how to integrate AI into work processes decision-making or problem-solving.
12 years of experience in software or platform engineering with 5 years in engineering management leading teams that own production platform services with a Bachelors degree; or 10 years and a Masters degree; or a PhD with 7 years of experience in Computer Science Engineering or a related technical field; or equivalent experience.
Proven track record managing teams that operate and scale open-source data infrastructure (query engines BI platforms developer environments or similar) in production.
Hands-on experience operating distributed query engines (Trino Presto Spark or similar) including cluster tuning scaling and performance optimization.
Strong knowledge of Kubernetes and containerized service deployment enough to architect solutions and debug issues even if a separate team owns the clusters.
Demonstrated ability to establish SLOs observability and incident-response practices for platform services and to drive operational maturity over time.
Experience managing platform upgrades migrations and version lifecycle for open-source technologies in production without disrupting users.
Proven people leadership. Experience hiring developing and retaining strong platform engineers and building team culture around operational excellence and automation.
Strong bias toward automation over manual toil with experience building or directing the development of internal tooling and self-service workflows.
Excellent collaboration skills across engineering data DevOps and business stakeholders.
Full professional proficiency in English.
Technical Expertise
Distributed query engines. Trino or Presto operations including deployment scaling resource group management query optimization connector configuration and upgrades.
Data catalog and lakehouse. Hive Metastore operations and familiarity with modern catalog alternatives (Nessie AWS Glue Unity Catalog Polaris). Apache Iceberg table format concepts.
BI and analytics platforms. Operating self-hosted BI tools such as Lightdash Redash Metabase or Superset including deployment scaling SSO integration and user management.
Developer platforms. Coder JupyterHub or similar cloud development environment platforms including workspace provisioning template management and resource policies.
Observability. Monitoring alerting and logging for platform services (Splunk Prometheus Grafana CloudWatch or similar). SLO design and tracking.
Security and access control. SSO/OIDC integration RBAC row-level security secrets management and audit logging across platform services.
Infrastructure familiarity. Kubernetes Helm Docker Infrastructure as Code (Terraform CDK) and CI/CD pipelines. Enough depth to partner effectively with infra teams and architect platform deployments.
Scripting and automation. Python Bash or Go for operational tooling automation and integration work.
Leadership & Communication
Proven ability to balance hands-on technical work with people leadership knowing when to go deep and when to delegate.
Strong technical judgment with the ability to evaluate open-source technologies make build-vs-buy decisions and sequence a platform roadmap.
Effective stakeholder management across technical and non-technical audiences translating platform capabilities and constraints into business terms.
Strong technical writing and documentation skills for runbooks architecture decisions and team processes.
Track record of building high-trust high-ownership engineering teams.
Nice to have
Direct experience operating Lightdash or dbt-integrated BI platforms.
Experience with Project Nessie or other versioned/transactional catalog systems.
Experience operating Coder or similar remote development environment platforms at scale.
Background in FinOps cloud cost management or financial data platforms.
Experience with Apache Iceberg table maintenance (compaction snapshot expiry partition evolution).
Experience in regulated or multi-environment cloud deployments (FedRAMP GovCloud or similar).
Open-source contributions to data infrastructure tooling.
Why join us
Own the platform services that power FinOps analytics for all of ServiceNows cloud spend at global scale.
Lead a team building on a modern fully open-source stack with real architectural influence.
Collaborate in a culture that values craftsmanship quality and innovation.
Work symbiotically with AI and automation tools that enhance engineering excellence and drive platform reliability.
Be part of a culture that encourages innovation continuous learning and shared success.
For positions in this location we offer a base pay of $190900 - $334100 plus equity (when applicable) variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline and individual total compensation will vary based on factors such as qualifications skill level competencies and work location. We also offer health plans including flexible spending accounts a 401(k) Plan with company match ESPP matching donations a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.
Additional Information :
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible remote or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation national origin age disability gender identity veteran status or any other category protected by addition all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process or are unable to use this online application and need an alternative method to apply please contact for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations including the U.S. Export Administration Regulations (EAR) ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. 2026 Fortune Media IP Limited. All rights reserved. Used under license.
Remote Work :
Yes
Employment Type :
Full-time
About Company
Learn here. Grow here. Make a difference here. At ServiceNow, our cloud?based platform and solutions deliver digital workflows that create great experiences and unlock productivity for employees and enterprises. Were growing fast, innovating even faster, and making an impact on our c ... View more