SeniorStaff Software Engineer, Distributed Systems
Santa Clara County, CA - USA
Job Summary
The Atomic Machines robotics fleet has reached a level of maturity where its ready to bring the Matter Compiler online. Now is the time manufacturing software must be brought up to par to leverage this fleet into a true fab. Concretely:
- Architectural decisions cloud vs. on-prem how far the control layer should extend are currently being made in parallel by people with different mental models of the end state without a shared written-down source of truth.
- The cost is real: engineering time is being spent today on work that may be thrown away once the architecture actually converges.
Were hiring a seasoned engineer who can operate productively inside that ambiguity someone who will write the design docs drive the team toward a decision and also model (and insist on) the testing review and traceability discipline that keeps this from happening again. This is as much a mandate to bring engineering rigor to the org as it is to build a platform that lets process developers and eventually in-the-loop physical AI safely program the fab. If that sounds like more org-building than you want in a software engineer role this probably isnt the right fit and thats a useful thing to know before either of us invests time in the process.
- Design and build the distributed software systems that coordinate state timing and behavior across manufacturing hardware; write test and debug sensors actuators and process controllers under real-time and reliability constraints.
- Design workflows that bridge manual and automated process steps and coordinate handoffs between production and process development.
- Evolve the existing system in place including but not limited to API development to govern machine behavior across our fleet as the architecture converges without waiting for a clean-slate rewrite to start delivering value.
- Instrument systems so machine and process data is legible not just to humans via logs but structured for downstream AI/ML consumption.
- Investigate and resolve issues that span software firmware and physical systems.
- Establish and model software engineering practices testing code review CI/CD documentation appropriate for a team thats outgrown its current ones.
- Contribute to system reliability through structured observability fault handling and graceful degradation.
- Collaborate closely with mechanical electrical and process engineers to translate physical constraints into resilient software behavior.
- Partner with engineering leadership to help converge competing architectural visions into one well-reasoned direction rather than waiting for it to be handed down.
- 5 years building or debugging systems with real external dependencies: hardware embedded devices networked services or similar.
- A track record of making and defending nontrivial architecture decisions build vs. buy deployment topology service boundaries not just implementing someone elses design.
- Strong Python skills for production systems plus proficiency in at least one systems or strongly-typed language (C Rust or Go) confirm against actual stack.
- Solid grounding in distributed systems fundamentals: state coordination consistency failure modes concurrency.
- Experience introducing or maintaining CI/CD automated testing or observability tooling in a codebase that didnt already have it.
- Comfort operating against an incomplete or contested spec and a bias toward driving clarity rather than waiting for it.
- Bachelors degree in Computer Science Electrical Engineering Robotics or related field or equivalent experience.
- Experience with real-time or resource-constrained environments or analogous domains (IoT edge compute industrial systems robotics warehouses manufacturing lines fabrication and automation facilities).
- Youve thought seriously about observability: what to instrument when logs arent enough how to make failure legible and youve instrumented systems specifically to make their data usable by ML/AI pipelines.
- Youve been the person who introduced testing or CI discipline to a team that didnt have it and can speak concretely about how you got buy-in not just what you built.
- Experience evaluating cloud vs. on-prem/edge deployment tradeoffs for latency- or safety-sensitive systems.
- Youve debugged issues that required reasoning across multiple system layers: application logic transport firmware hardware.
- Youre genuinely energized by translating physical constraints latency noise mechanical tolerance safety margins into software behavior.
The compensation for this position also includes equity and benefits.
Salary Range
$180000 - $230000 USD
Required Experience:
Staff IC