Site Reliability Engineer Grid Orchestration Software
Bellevue, WA - USA
Job Summary
The ideal candidate operates with a high degree of autonomy applying sound technical judgment within established policies and procedures. This individual will play a key role in software implementation troubleshooting customization and integration within customer environments while effectively balancing project scope timelines and resource commitments.
Roles and Responsibilities
- System Integration: Integrate Grid Orchestration Software (GridOS) solutions and perform comprehensive system testing.
- Incident Management: Lead cross-functional incident response to resolve production issues and minimize customer impact.
- Post-Incident Analysis: Conduct reviews to drive corrective actions that decrease outage frequency and severity.
- Monitoring & Observability: Design and maintain monitoring logging and alerting platforms (e.g. Prometheus Grafana).
- Engineering Partnership: Collaborate with software teams to improve system resilience release quality and operational readiness.
- Reliability & Planning: Drive continuous improvement in disaster recovery capacity planning and system architecture design.
- Automation: Develop scripts and workflows to streamline operations and reduce manual effort.
- Leadership & Support: Mentor team members and serve as a technical resource for the organization.
- Client Relations: Resolve customer technical issues provide updates and manage on-call responsibilities as required.
- Travel includes less than 5% typically but depending on the project you are involved in it could be more
- We will not provide Visa sponsorship now or in the future for this role.
- Relocation assistance is available for U.S. based candidates only.
Required Qualifications:
- Bachelors degree from an accredited university or college (or a high school diploma / GED with at least 6 years of experience in Job Family Group(s)/Function(s)).
- 3 years of experience in site reliability engineering DevOps systems engineering or software engineering
- 3 Proficient with containerization and orchestration tools such as Docker and Kubernetes
Desired Characteristics:
- Strong problem-solving skills with the ability to perform effectively under pressure
- Excellent cross-functional communication and collaboration skills
- Deep familiarity with Linux/Unix system administration and networking fundamentals (TCP/IP DNS HTTP).
- Experience with monitoring and observability tools such as Prometheus Grafana or ELK
- Strong oral and written communication skills.
- Experience operating and supporting large-scale distributed systems
- Familiarity with CI/CD tools such as Jenkins GitHub Actions GitLab CI or Azure DevOps
- Familiarity with database reliability backup and recovery processes
GE Vernova offers a great work environment professional development challenging careers and competitive compensation. GE Vernova is anEqual Opportunity Employer. Employment decisions are made without regard to race color religion national or ethnic origin sex sexual orientation gender identity or expression age disability protected veteran status or other characteristics protected by law.
GE Vernova will only employ those who are legally authorized to work in the United States for this opening. Any offer of employment is conditioned upon the successful completion of a drug screen (as applicable).
Relocation Assistance Provided: No
Required Experience:
IC
About Company
GE Vernova's Asset Performance Management software can help you increase asset reliability, minimize costs and reduce operational risks. View a demo today.