Enterprise Platform Architect
Palo Alto, CA - USA
Job Summary
About Client
They are transforming core enterprise functions by combining frontier AI strong engineering and deep domain expertise. Our platform is used by a growing number of Fortune 1000 companies to execute complex high-value work. We are building enterprise AI for environments where accuracy reliability security scalability and economics matter.
They are looking for an experienced professional to further push the availability reliability scalability performance and cost efficiency of our platform. You will work across Engineering Cloud Infrastructure DevOps QA and AI teams to establish measurable requirements identify architectural platform bottlenecks and opportunities and define enhancements.
- Availability & Resilience
- Define measurable availability resilience and recovery requirements.
- Architect for failures across infrastructure APIs databases external dependencies and AI models.
- Design redundancy failover retries timeouts graceful degradation and recovery mechanisms.
- Identify and eliminate critical single points of failure.
- Establish resilience failover and recovery testing with clear production-readiness metrics.
- Identify gaps drive remediation and validate readiness for enterprise production.
- Scalability & Performance
- Define measurable targets for throughput concurrency latency document size storage growth and model capacity.
- Architect the platform to scale predictably across customers workloads and data volumes.
- Lead capacity planning across compute storage databases networking and AI infrastructure.
- Establish load stress endurance and performance-testing standards.
- Identify and eliminate architectural and performance bottlenecks.
- Maintain performance benchmarks and ensure the platform meets enterprise-scale requirements before production.
- Cost Efficiency
- Define and track platform unit economics including cost per transaction workflow and AI execution.
- Establish cost targets and identify the primary drivers of platform economics.
- Optimize model selection routing caching batching and reuse.
- Move workloads from expensive LLM reasoning to code ML smaller models or deterministic systems where appropriate.
- Improve infrastructure utilization and eliminate unnecessary computation.
- Ensure the platform remains economically viable as workload volume and complexity scale.
- Experience: 10 years of experience architecting and building enterprise SaaS platforms or similarly complex production systems. Experience in enterprise tech companies is highly desired.
- Education: Bachelors or Masters degree in Computer Science Engineering or a related field.
- Technical Skills:
- Deep expertise in distributed systems reliability scalability performance and cloud architecture.
- Strong experience with Azure or other hyperscale cloud platforms.
- Strong understanding of databases APIs networking storage containers and distributed compute.
- Familiarity with AI/LLM infrastructure model APIs inference architectures and AI economics.
- Experience with observability load testing capacity planning resilience engineering and disaster recovery.
- Ability to make sound architectural tradeoffs across reliability performance complexity and cost.
- Ability to lead architecture and drive execution across multiple engineering teams.
- Competitive Compensation: Tailored to your experience and skill set.
- Flexible Work Arrangements: Hybrid working model for work-life balance.
- Career Growth: Opportunities for professional development and leadership roles.
- Innovative Culture: Work on transformative technologies and make an impact in the AI space.
Performance Bonus: 10 - 15% (Distributed Quarterly)
Equity Options: Eligible
PTO: Unlimited
Required Experience:
Staff IC