Enter a job title or keyword

AI DevOps Engineer


Job Location:

Pune - India

Monthly Salary: Not provided by the employer
Experience Required: 5-6years
Posted: 30 September 2026 (9 hours ago)
Application Deadline: 28 December 2026
Vacancies: 1 Vacancy

Job Summary

About the role :-

We are hiring an AI DevOps Engineer to own how our AI systems get deployed stay up and stay observable. You will own the infrastructure and operational substrate for nCircle Techs agentic platform the CI/CD pipelines the cloud environments the release machinery the reliability practice and the telemetry layer that turns opaque agent behavior into something you can measure alert on and debug.

Our engineers build the AI agents; you make shipping them safe repeatable and observable. Right now most of our apps have no CI/CD the infrastructure is fragmented and the agentic systems being stood up have little more than print statements for observability. Building that operational foundation is the job.

Agentic systems fail differently from ordinary services nondeterministic output silent quality drift runaway tool-call loops and cost that spikes without warning. Standard DevOps is necessary but not sufficient. This role exists because someone has to own reliability and telemetry for systems that dont fail the way the runbooks assume.

What youll do :-

Own CI/CD for the platform. Build the pipelines that take AI systems from commit to production automated testing evaluation gates security and dependency checks controlled and canary releases and one-command rollback. Most of our apps have no pipeline today; you will establish the pattern and make the safe path the default path.

Manage cloud infrastructure as code. Own the cloud environments (AWS primarily) the platform runs on provisioning networking secrets environment parity and cost controls as versioned reviewable infrastructure-as-code not hand-tuned consoles. You are accountable for environments that are reproducible least-privilege by default and cheap to stand up and tear down.

Run the reliability practice. Own production reliability: SLOs on-call and incident response capacity and cost management self-healing loops that detect and recover from failures and blameless post-incident review. You will help define what up and healthy even mean for a nondeterministic system.

Build the agent telemetry and observability layer. Instrument the platform so agent behavior is legible: structured traces of agent runs and tool calls token and cost accounting latency and success metrics output-quality tracking over time and the dashboards and alerts that surface a regression before a user does. When an agent misbehaves in production the telemetry you built is how the team finds out and figures out why.

Set the operational standard by example. On a small high-leverage team your pipelines and dashboards are the template. You establish the deployment patterns others adopt the observability every new system gets wired into by default and the operational discipline that lets a lean team run production systems well.

What were looking for :-
  • 68 years in DevOps SRE platform or infrastructure engineering with a track record of running production systems you were accountable for
  • Deep CI/CD experience you have built and owned pipelines (GitHub Actions GitLab CI or similar) that gate test and safely release real production software
  • Strong cloud operations ideally AWS provisioning networking secrets and cost management as infrastructure-as-code (Terraform CDK or similar)
  • Hands-on observability experience metrics logging distributed tracing dashboards and alerting (OpenTelemetry Prometheus/Grafana Datadog CloudWatch or similar) and the instinct to instrument first
  • SRE fundamentals: SLOs incident response on-call capacity planning and blameless postmortems
  • Enough software fluency to read application code wire telemetry into it and debug a failing deploy without waiting for someone else
Nice to have :-
  • Experience operating AI or LLM systems in production token and cost accounting prompt and evaluation-score tracking or LLM observability tooling (LangSmith Langfuse Arize or similar)
  • Familiarity with the failure modes of nondeterministic systems: quality drift runaway loops cost spikes non-reproducible output
  • Experience with Databricks or a similar lakehouse platform and with tool-integration layers such as MCP
  • Container and orchestration experience (Docker Kubernetes or serverless equivalents)
  • Experience in a non-software-company engineering organization internal tools corporate IT transformation or similar
  • Experience standing up an internal platform or golden-path deployment pattern that other teams adopted
Why this role :-

You will build the operational foundation an entire organizations AI runs on the pipelines the environments and the telemetry with a clear mandate and a direct line to the Director of APEX. The platform is early and that is the appeal. You are not tuning someone elses mature platform; you are building the deployment and observability substrate nCircle Tech will run AI on for the next decade and defining what running AI in production looks like here.