Site Reliability Engineer
Chicago, IL - USA
Job Summary
Nextpoint builds transformative software and services for the legal industry making eDiscovery case management and litigation prep simple fluid and affordable for law firms of all sizes. Our secure cloud-based platform lets teams start document review in minutes backed by powerful analytics an intuitive interface and best-in-class security at every point.
Were problem solvers simplifiers and challenge seekers united by a shared goal: a great team culture and satisfied clients. Were headquartered in Chicagos Ravenswood neighborhood and proud to have been named one of Built Ins Best Startups to Work for in Chicago five years running ().
This role is hybrid (at least 3 days/week) in our Chicago office. Nextpoints platform handles massive unpredictable volumes of sensitive legal data unlimited-upload document review AI-assisted analysis and secure electronic production all running on AWS with zero downtime tolerance for firms in active litigation. Were looking for a Site Reliability Engineer to help maintain the reliability scalability and security posture of that platform as we expand our AI capabilities (built on Amazon Bedrock) and grow our customer base.
This is a hands-on role for someone who wants to be the first line of response for production infrastructure at a company where reliability is a customer-trust issue not just an engineering metric.
RESPONSIBILITIES
Infrastructure & Automation (50%)
Maintain and extend existing infrastructure-as-code (Terraform/CloudFormation/CDK) following established patterns and standards
Support and operate CI/CD pipelines; implement improvements as directed
Monitor cloud cost trends and flag optimization opportunities for review
Reliability & Operations (25%)
Monitor uptime latency and performance SLOs/SLIs for production systems supporting document upload processing review and production workflows
Participate in the on-call rotation and serve as first responder for production incidents during US business hours
Triage troubleshoot and resolve incoming infrastructure requests and incidents; escalate and coordinate on complex root-cause work
Write clear post-incident reports and help investigate recurring incidents and cost overruns
Maintain and extend existing monitoring alerting and observability tooling
Cross-Functional Collaboration (15%)
Partner with engineering teams to build reliability scalability and observability into new features from design through launch
Document runbooks architecture decisions and operational procedures for the broader engineering team
Communicate incident status and technical issues clearly to engineering and non-technical stakeholders
Participate in design reviews to flag reliability or operational concerns early
Security & Compliance (10%)
Support SOC 2 compliance activities and help uphold encryption access-control and audit-trail standards across all environments
Implement and maintain security best practices for infrastructure handling confidential legal and client data including AI workloads on Amazon Bedrock
Support security reviews vulnerability management and patching cadences across production systems
Maintain permissions-based access controls and comprehensive audit logging in line with client security commitments
5 years of experience in Site Reliability Engineering DevOps or Infrastructure Engineering roles ideally in a B2B SaaS environment
Deep hands-on experience with AWS (EC2 S3 RDS Lambda VPC IAM CloudWatch or equivalent services)
Working knowledge of infrastructure-as-code (Terraform CloudFormation or CDK) and configuration management
Proficiency in at least one scripting/programming language (Python Go or similar) for automation and tooling
Experience with containerization and orchestration (Docker Kubernetes or ECS)
Track record of operating CI/CD pipelines
Experience with monitoring/observability stacks (Datadog CloudWatch Prometheus/Grafana or similar)
Familiarity with security compliance frameworks (SOC 2 HIPAA or similar) and encryption/access-control best practices
Experience supporting systems handling large-scale variable-volume data processing is a plus
Exposure to AI/ML infrastructure (e.g. Amazon Bedrock model-serving pipelines) is a plus given our growing AI feature set
Experience using ClaudeCode Kiro OpenCode or similar Agentic AI
Bachelors degree in Computer Science Engineering or related field (or equivalent experience)
Strong written and verbal communication skills with the ability to write clear runbooks and explain technical tradeoffs to non-technical stakeholders
Comfortable being part of an on-call rotation
EQUAL OPPORTUNITY EMPLOYER
Nextpoint is an equal opportunity employer. We actively work to build a diverse team and encourage candidates of all backgrounds to apply. All applicants are considered without regard to race color religion sex sexual orientation gender identity national origin veteran status disability or any other characteristic protected by applicable law.
Required Experience:
IC
About Company
Hey, thanks so much for joining us! We totally get how exhausting job hunting can be, so let’s dive right in and share what we’ve got for you. Who We are and wh...