DataOps Engineer
Job Location:
Englewood Cliffs, NJ - USA
Monthly Salary:
Not Disclosed
Posted on:
3 hours ago
Vacancies:
1 Vacancy
Job Summary
- Iceberg operations: support tables manage schema changes partitions snapshot retention and keep the catalog (Hive Metastore AWSGlue Nessie ) synchronized.
- Docker image creation & testing: write multistage Dockerfiles for Spark/Flink/Presto run local test environments with DockerCompose and conduct vulnerability scans (Trivy Snyk ).
- Data pipeline development: build ETL/ELT jobs that ingest raw data and write to Iceberg tables; add simple streaming components using Kafka Pulsar or Kinesis when needed.
- CI/CD automation: configure pipelines (GitHubActions GitLabCI AzureDevOps ) to lint Dockerfiles scan images version Iceberg metadata and deploy pipelines without downtime.
- Automation with Ansible/Python: script cluster provisioning catalog configuration vacuum/compaction and other routine housekeeping tasks.
- Observability: instrument services with OpenTelemetry Prometheus Grafana and Loki; create dashboards showing pipeline latency resource usage table health and error rates; set up basic alerts.
- SLA monitoring: measure data freshness job success rates and query response times against agreedupon targets and report deviations.
- Incident response: join the oncall rotation perform firstline diagnosis and resolution of pipeline failures Iceberg metadata issues or container crashes; write concise rootcause analyses and suggest improvements.
- Security & compliance support: help enforce image signing mTLS IAM roles and bucket policies; collaborate with the security team to meet GDPR HIPAA or ISO27001 requirements.
- Knowledge sharing: keep internal documentation up to date and run short tech demos or brownbag sessions on Iceberg Docker best practices and automation techniques.
Qualifications :
Requirements
- Bachelors degree in Computer Science IT Data Engineering or a related field (Masters a plus).
- 5years of handson experience building and operating largescale data platforms (lakehouse datawarehouse or bigdata ecosystems).
- Proven production experience with Apache Iceberg (table creation partition management schema evolution catalog integration).
- Strong Docker skills: multistage builds DockerCompose testing routine image security scanning.
- Experience with at least one major dataprocessing engine (Spark Flink or Presto/Trino) and its connection to Iceberg tables.
- Proficiency in Python and/or Ansible for automating infrastructure and platform tasks.
- Experience building CI/CD pipelines that include Docker linting vulnerability scanning and automated deployment of datapipeline code.
- Familiarity with observability tooling (Prometheus Grafana OpenTelemetry Loki) and ability to create useful alerts and dashboards.
- Ability to respond to incidents write clear rootcause analysis reports and contribute to postmortem actions.
- Willingness to participate in an oncall rotation as a firstline responder.
- Availability to work onsite in NewJersey for the initial assignment and relocate to Dallas by October2026.
Preferred Qualifications
- Experience with cloudnative data services on AWS Azure or GCP (EMR Dataproc Synapse etc.).
- Familiarity with other lakehouse formats such as DeltaLake or ApacheHudi and ability to evaluate tradeoffs against Iceberg.
- Knowledge of streaming platforms (Kafka Pulsar Kinesis) and realtime processing patterns.
- Relevant certifications (Databricks Lakehouse Associate Google Professional Data Engineer AWS Certified Data Analytics Specialty etc.).
- Background supporting data platforms in regulated industries (pharma finance healthcare) and understanding of associated compliance frameworks.
Additional Information :
All your information will be kept confidential according to EEO guidelines.
Remote Work :
No
Employment Type :
Contract