Software Development Engineer MultiCloud
Austin, TX - USA
Job Summary
Software Development: Design develop test and maintain software services APIs tools and automation using languages such as Python and u0026 Platform Engineering: Build and improve cloud-native services and platform capabilities across production and non-production : Deploy operate and troubleshoot applications and services running on Kubernetes. nDevelop automation that simplifies deployment and platform : Apply reliability engineering practices to improve system availability performance scalability and operational : Identify repetitive operational activities and develop software and automation to reduce manual effort and operational : Use logs metrics traces dashboards and alerts to understand system behavior troubleshoot issues and identify opportunities for : Analyze operational and application data to identify trends anomalies recurring issues and performance bottlenecks. Develop dashboards and reporting that provide actionable Infrastructure: Support infrastructure and platform capabilities for AI/ML training LLM inference and GPU-based workloads. Develop automation to simplify deployment and operation of these -Assisted Engineering: Explore and apply LLMs and AI technologies to improve software development troubleshooting analytics automation and operational u0026 Resilience: Participate in capacity planning performance testing scale testing and disaster recovery Improvement: Identify opportunities to improve platform reliability developer experience automation and operational : Create and maintain technical documentation operational procedures troubleshooting guides and : Work closely with software engineering platform SRE QA AI/ML security architecture and program management teams.
Software Engineering: Solid understanding of software engineering fundamentals and experience developing and maintaining production software services tools or : Proficiency in Python and/or Go (Golang) with the ability to write clean maintainable and testable : Hands-on experience with Kubernetes and containers including deploying and troubleshooting applications. Familiarity with Helm Kustomize or similar tools is : Experience with at least one major cloud platform such as AWS Google Cloud or as Code: Familiarity with Terraform Ansible or similar infrastructure automation Engineering: Understanding of reliability concepts such as monitoring alerting SLIs/SLOs incident management capacity planning and : Experience with or exposure to technologies such as Prometheus Grafana Splunk OpenTelemetry or similar observability : Ability to analyze system and application data to identify trends and troubleshoot issues. Familiarity with SQL Python-based data analysis dashboards or reporting tools is a Systems: Working knowledge of distributed system concepts including availability scalability networking fault tolerance and performance.
AI/ML: Familiarity with AI/ML concepts LLMs model inference or GPU workloads is a plus. Prior AI/ML infrastructure experience is beneficial but not Solving: Strong analytical and troubleshooting skills with an interest in solving problems through software and : Strong communication skills and the ability to work effectively within cross-functional engineering teams.
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more