Senior Site Reliability Engineer (Data Platform)
Kuala Lumpur - Malaysia
Job Summary
Summary
The Team and the OpportunityYou will join the PDO Site Reliability Engineering (Data Platform) team that owns the reliability and operability of Guidewires data platform services including largescale data processing analytics and streaming capabilities that underpin our AI and Insight products.
This team partners closely with product engineering data platform and security to design for reliability build automation and run services in production.
As a Senior Site Reliability Engineer (Data Platform) you will be a technical leader in running and evolving our big data stack on AWS (and potentially other public clouds) using software engineering to solve infrastructure and application reliability challenges. Youll help advance
PDOs priorities by improving incident response hardening critical data paths and enabling scalable costefficient operations for our customers most important workloads.
Job Description
What You Will Do
Design and implement selfservice automation and tooling (in Go Python or scripting languages) to standardize and streamline deployment operations and troubleshooting for data platform services.
Implement and improve CI/CD pipelines (e.g. TeamCity Github Actions) to support safe frequent deployments including gate promotion and automated quality checks.
Use Infrastructure as Code (e.g. Terraform AWS CloudFormation) to build harden and maintain repeatable cloud infrastructure for data and analytics workloads.
Operate and improve Kubernetesbased environments (AWS EKS) including deployment scaling and lifecycle management of containerized data services (e.g. Dockerpackaged microservices streaming jobs).
Apply progressive delivery strategies such as blue/green and canary deployments and support chaos engineering experiments to validate resilience and recovery mechanisms.
Collaborate on capacity planning and costaware design for cloud resources across compute storage and networking layers for dataintensive systems.
Build and refine endtoend observability for the data platform using monitoring and logging tools (e.g. Datadog ELK) including metrics traces and logs.
Develop meaningful dashboards and alerts to provide clear visibility into data pipeline health customer experience and platform performance.
Analyze operational data to identify reliability risks and bottlenecks feeding insights into the roadmap and reliability backlogs.
Partner with product engineering data platform security and other SRE teams to define and implement improvements in service architecture and operational practices that support PDOs AI cloud and data platform priorities.
Advocate for reliability resilience and operational excellence in design reviews readiness assessments and release planning.
Contribute to a positive inclusive work environment based on accountability continuous learning and psychological safety consistent with Guidewires culture of determination collaboration continuous improvement and bravery.
What You Need to Succeed
Experience and Education
8 years of relevant industry experience in Site Reliability Engineering DevOps Production Engineering or similar roles supporting largescale distributed systems and data platforms.
BS/MS in Computer Science Computer Engineering Mathematics or equivalent practical experience.
Technical Skills
Strong experience with continuous deployment and operation of cloud services on public cloud (AWS) including production support and oncall.
Hands-on experience running data platforms using big data and streaming technologies such as Kafka Hadoop Spark and Hive on the public cloud.
Proficiency in at least one of Java Go or Python and solid skills with scripting languages to build tools automation and integrations.
Experience building and operating microservices including REST APIs and/or gRPC services.
Solid experience with CI/CD tools (e.g. TeamCity Github Actions) for automated builds tests and deployments including promotion gates.
Strong experience with Infrastructure as Code tools such as Terraform and AWS CloudFormation for provisioning and managing cloud infrastructure. Familiarity with Kubevela/Crossplane is a plus
Practical knowledge of Kubernetes (e.g. AWS EKS) and Docker including deployment patterns service discovery and resource management.
Familiarity with AWS services relevant to data and distributed systems such as RDS EMR Redshift MSK (Managed Streaming for Kafka) ECS SNS and SQS. Expertise with monitoring logging and observability tools (e.g. Datadog ELK) to instrument services and build actionable alerts and dashboards.
Deep understanding of distributed systems fundamentals networking storage operating systems and how they interact in complex multitier environments.
Knowledge of capacity planning scalability and resilience patterns (including blue/green and canary deployments and chaos engineering concepts).
Operational and ProblemSolving Skills
Demonstrated experience solving infrastructure and application problems using software engineering approaches rather than only manual operations.
Familiarity with agile methodologies like Scrum and Kanban.
Strong analytical and troubleshooting skills for complex distributed multiservice environments.
Experience with oncall incident response (e.g. PagerDuty) and postincident review processes with a bias for learning and continuous improvement.
Ways of Working
Ability to collaborate effectively with other engineering data and operations teams to understand their systems and help improve them.
A bigpicture perspective on systems tools and customer value aligning technical decisions with PDOs priorities around operational excellence AI cloud and data platform adoption.
Comfort with agile development methodologies and iterative delivery in a highly collaborative environment.
Eagerness to learn experiment and growstaying current with emerging technologies across cloud data and SRE practices and applying them thoughtfully where they add real value.
Bonus Points
Kubernetes/AWS certifications
Contributions to open source projects
#LI-AA1
About Guidewire
Guidewire is the platform P&C insurers trust to engage innovate and grow efficiently. We combine digital core analytics and AI to deliver our platform as a cloud service. More than 540 insurers in 40 countries from new ventures to the largest and most complex in the world run on Guidewire.
As a partner to our customers we continually evolve to enable their success. We are proud of our unparalleled implementation track record with 1600 successful projects supported by the largest R&D team and partner ecosystem in the industry. Our Marketplace provides hundreds of applications that accelerate integration localization and innovation.
For more information please visit and follow us on Twitter: @GuidewirePandC.
Guidewire Software Inc. is proud to be an equal opportunity and affirmative action employer. We are committed to an inclusive workplace and believe that a diversity of perspectives abilities and cultures is a key to our success. Qualified applicants will receive consideration without regard to race color ancestry religion sex national origin citizenship marital status age sexual orientation gender identity gender expression veteran status or disability. All offers are contingent upon passing a criminal history and other background checks where its applicable to the position.
Required Experience:
Senior IC
About Company
Elevate your P&C insurance with Guidewire's industry-leading software! Streamline workflows, enhance customer experience, and drive growth. Learn more today!