Senior DevOpsPlatform Engineer
Buenos Aires - Argentina
Job Summary
Part-Time Contract Remote Approximately 20 Hours per Week
About the Role
A mission-critical SaaS platform used by approximately 800 flooring businesses to manage leads estimates product selection inventory scheduling installations invoicing and payments is being modernized from a /ASP SQL Server stored-procedure-heavy environment to React TypeScript native PostgreSQL REST APIs and AWS. Much of the initial conversion has been completed using AI-assisted development. The immediate priority is to clean up test validate and safely release the converted system while existing customers remain active.
We are looking for a hands-on Senior DevOps/Platform Engineer who can strengthen the delivery platform observability security and production safeguards required for frequent low-risk releases.
Responsibilities
- Assess the current AWS infrastructure CI/CD pipelines environments deployment process and production risks.
- Improve automated build test security-scanning deployment and rollback workflows.
- Support small frequent releases through feature flags staged rollouts health checks and automated validation.
- Establish safe coexistence between legacy and modern services during the migration.
- Build and maintain infrastructure as code for repeatable auditable environment provisioning.
- Strengthen monitoring centralized logging distributed tracing alerting dashboards and incident-response procedures.
- Define service-level indicators for availability latency error rates database health queues and critical business workflows.
- Help diagnose performance and stability issues involving AWS services APIs PostgreSQL Babelfish and components.
- Support the transition from Babelfish and SQL Server-oriented behavior to native PostgreSQL.
- Review capacity connection-pool timeout retry queue and load-shedding configurations.
- Improve resilience for asynchronous and event-driven workloads using services such as Lambda SQS EventBridge API Gateway Redis and dead-letter queues.
- Implement secure secrets management access controls audit logging vulnerability management backup and recovery procedures.
- Verify that rollback remains safe after application or database writes including backward-compatible schema changes.
- Review AI-generated infrastructure and deployment changes for correctness security maintainability and operational risk.
- Document deployment standards incident procedures environment configuration and operational ownership.
- Partner closely with engineering and product teams with meaningful overlap during US working hours.
Required Qualifications
- At least 7 years of professional DevOps SRE platform engineering or cloud infrastructure experience.
- Strong hands-on experience operating production workloads in AWS.
- Proven ownership of CI/CD architecture automated deployments monitoring incident response and rollback.
- Experience supporting production systems built with React PostgreSQL .
- Strong understanding of:
- Infrastructure as code
- Networking DNS load balancing and API gateways
- Containers and serverless workloads
- Logs metrics traces alerts and production dashboards
- Secrets management and least-privilege access
- Backups disaster recovery and production readiness
- Database migrations and backward-compatible schema changes
- Connection pooling capacity planning retries idempotency and failure recovery
- Experience introducing staged deployments feature flags canary releases or blue-green deployments.
- Ability to troubleshoot across application infrastructure network and database layers.
- Experience working on incremental modernization where legacy and modern systems must run in parallel.
- Strong written and spoken English with the ability to communicate risk and recommendations clearly.
- Ability to work independently and deliver senior-level ownership within a part-time schedule.
Preferred Qualifications
- Experience with AWS Lambda SQS EventBridge API Gateway Redis CloudWatch and container-based AWS services.
- Experience migrating SQL Server workloads to PostgreSQL particularly through Babelfish.
- Experience with Terraform CloudFormation or similar infrastructure-as-code tooling.
- Experience operating workflow-heavy SaaS involving inventory scheduling invoicing payments or other production-critical transactions.
- Familiarity with OpenTelemetry distributed tracing and structured logging.
- Experience reviewing infrastructure or deployment code generated through Claude Code Codex Cursor or similar AI tools.
- Experience improving engineering platforms for teams releasing daily or near-daily.
What Success Looks Like
During the initial engagement this person will:
- Produce a clear assessment of the current infrastructure deployment risks and release blockers.
- Establish reliable CI/CD and environment-management standards.
- Improve monitoring across the core quote-to-invoice workflow.
- Make deployments smaller repeatable observable and safely reversible.
- Strengthen production readiness for the AI-converted React and PostgreSQL platform.
- Reduce instability associated with Babelfish database connections asynchronous workloads and legacy-modern coexistence.
- Create practical operational documentation the engineering team can maintain after the engagement.
Engagement Details
- Remote part-time contract
- Approximately 20 hours per week
- Regular overlap with the US-based team required
- Initial focus on release readiness and platform stabilization followed by modernization scalability and support for AI-enabled product capabilities