Software Engineer, Ontology
San Francisco, CA - USA
Job Summary
We exist to make humanity more free. For most of human history you farmed or you starved. Technology gave people more time for the things they wanted to do instead of things they had to do. Powerful AI will be the biggest lever for human choice weve ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who dont share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.
Were singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else rethinking every layer of the stack. We acquire power design and build data centers and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.
We hire people who care deeply about this problem space. If that is you please apply!
Be a barrel. Full autonomy. Own things end to end take on scope without being asked no permission required to operate outside your core role.
Insane urgency. We drive everything forward as fast as possible.
Reason from first principles. Challenge every assumption. Zero analogy thinking no egos the best idea wins.
Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
Build something that actually matters. If youre going to spend your time spend it on something that matters to the world.
Fluidstack a leading cloud provider is looking for a Platform Engineer to build the foundational platforms that enable our global infrastructure and data center operations. Youll develop comprehensive internal tooling across multiple domainsCMDB asset management DCIM monitoring and observability security and operational automationthat streamline how we deploy manage and operate infrastructure at scale. Working cross-functionally with engineering operations data center teams and product youll deliver scalable reliable user-friendly solutions that directly impact our ability to grow and deliver world-class infrastructure services.
Infrastructure Platform Development
Design and build our next-generation CMDB system as the authoritative source of truth for infrastructure assets network topology and configuration data
Create DCIM platforms for rack operations server/GPU deployment OS installation quality assurance and white-screen operations
Develop end-to-end asset lifecycle management systems covering receiving racking inventory break-fix and decommissioning workflows
Build monitoring and observability platforms integrating telemetry from BMS EPMS and IT devices with intelligent alarming and incident management
Create self-service portals and automation for new region bootstrap day-2 operations and fleet-scale management
Operational Excellence & Automation
Eliminate manual toil through workflow automation and self-service tooling that empower operations and engineering teams
Build workflow orchestration systems for complex multi-step processes spanning incident problem and change management
Develop digital twin visualizations and operational dashboards surfacing actionable insights; partner with data teams on analytics
Create integration layers connecting internal platforms with external vendors and third-party systems
Cross-Functional Partnership
Collaborate with data center operations system engineering network engineering and security teams to understand requirements and deliver high-impact solutions
Work with product and business stakeholders to prioritize features define roadmaps and balance competing needs
Align with support and operations teams to ensure platforms scale with organizational growth
Technical Leadership
Evaluate build vs. buy decisions for platform components weighing in-house development against commercial SaaS and open-source solutions for scalability cost and flexibility
Champion modern development practices including CI/CD infrastructure-as-code automated testing and observability-first design
Participate in architecture reviews and design discussions contributing to technical direction and standards
Foster technical excellence through code reviews documentation and knowledge sharing
Scalability & Reliability
Design high-performance fault-tolerant systems capable of handling thousands of QPS as our infrastructure footprint expands
Build comprehensive monitoring logging and debugging capabilities with robust error handling
Implement data migration strategies and manage upstream/downstream dependencies carefully during platform evolution
Own projects end-to-end from concept through deployment ensuring production readiness and operational excellence
3 years of professional software development experience building production systems
Strong programming skills in Python Go or similar languages with understanding of system design patterns
Experience designing and implementing RESTful APIs data models and distributed systems
Proficiency with relational and NoSQL databases (PostgreSQL Redis etc.)
Hands-on experience with containerization (Docker) and infrastructure-as-code tools (Terraform Ansible)
Understanding of CI/CD pipelines and modern development workflows
Solid grasp of networking fundamentals (TCP/IP DNS HTTP) and Linux/Unix environments
Strong problem-solving abilities with attention to scalability reliability and operational concerns
Excellent communication skillsable to convey technical concepts to both technical and non-technical stakeholders
Experience with CMDB systems (NetBox Device42) or asset management platforms
Background in infrastructure automation DevOps or platform engineering
Familiarity with workflow orchestration frameworks (Temporal Airflow Camunda)
Knowledge of monitoring and observability stacks (Prometheus Grafana OpenTelemetry)
Experience with time-series databases and data visualization
Understanding of ITSM frameworks (ITIL) and service management practices
Experience in data center operations facilities management or physical infrastructure
Contributions to open-source infrastructure projects
Bachelors degree in Computer Science or equivalent practical experience
Competitive total compensation package (salary equity).
Retirement or pension plan in line with local norms.
Health dental and vision insurance.
Generous PTO policy in line with local norms.
Total compensation may also include equity in the form of restricted stock units.
We are committed to pay equity and transparency.
Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race color religion sex national origin sexual orientation gender identity disability and protected veterans status or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email please email with your resume/CV the role youve applied for and the date you submitted your application-- someone from our recruiting team will be in touch.
Required Experience:
IC