Mid-Level Data Engineer6+yrsHyderabad
Job Summary
IT Development Team
Overview
Location: onsite Hyderabad
Employment Type: full time
Experience: 6 years
Compensation: INR-
Key Skill Requirements:
- Strong experience in Python
- Hands-on expertise with AWS services particularly:
- EMR (mandatory and most critical skill)
- EC2
- Lambda
- Candidate should have experience with EMR performance and cost optimization
- Other AWS services can be flexible based on overall profile strength
Other notes:
Strong Real-World Data Engineering Experience
Not just theoretical knowledge.
Wants engineers who can demonstrate:
- Practical implementation experience
- Real production environments
- Solving actual business problems
- Experience working through technical challenges
- Understanding why decisions were made not just what was built
- Looking for people with practical knowledge and real-time experience.
Candidates Who:
- Are naturally curious
- Continuously learn
- Adapt quickly to new technologies
- Enjoy solving new problems
- Demonstrate initiative
- Show flexibility as technologies evolve
Experience Required
Extracted: 6 years of experience in data engineering distributed systems or backend platforms
Overview
Senior Data Engineer role focused on designing developing and optimizing large-scale data platforms and backend systems. The position emphasizes building API-driven data services managing distributed batch data processing and optimizing cloud workflows on AWS (with mandatory EMR). Work includes orchestration with Apache Airflow and building/maintaining pipelines using Spark SQL Hive Python and Scala.
Key Responsibilities
Design and develop API-driven systems for managing large-scale batch data applications
Build scalable backend services and data engineering solutions
Develop and maintain data pipelines using Spark SQL Hive Python and Scala
Design and optimize workflows using Apache Airflow (DAG design scheduling monitoring failure handling)
Work with AWS services including EMR EC2 S3 Lambda DynamoDB and API Gateway
Optimize EMR performance and cost efficiency for large-scale data workloads
Participate in architecture reviews code reviews and performance tuning initiatives
Collaborate with cross-functional teams including product QA DevOps and engineering
Troubleshoot complex production issues and improve system reliability
Support CI/CD processes automation and operational excellence initiatives
Contribute to system design testing and deployment strategies
Perform other duties as assigned
Required Qualifications
6 years of experience in data engineering distributed systems or backend platforms
Strong proficiency in Python or Scala for building production-grade pipelines
Hands-on experience with AWS cloud services (EMR is mandatory)
Strong experience with Spark SQL Hive and data processing frameworks
Advanced experience with Apache Airflow orchestration
Experience designing and supporting large-scale data pipelines and batch processing systems
Strong understanding of data engineering concepts (data transformation optimization data quality)
Experience with version control and CI/CD tools (Git Jenkins or equivalent)
Strong analytical troubleshooting and problem-solving skills
Excellent communication and collaboration skills
Technical Skills Required
AWS EMR
AWS
Apache Spark
SQL
Hive
Python
Scala
Apache Airflow
Git
Jenkins
CI/CD
Batch processing
Data pipelines
Distributed systems
Data transformation
Data quality
Troubleshooting
Nice-to-have
EMR performance tuning
Cost optimization
API design
Microservices
Event-driven architectures
AI/ML
LLMs
AI-assisted engineering
Data governance
Data lineage
Observability tools
Real-time data streaming architectures
Tools/Platforms
AWS EMR
AWS EC2
AWS S3
AWS Lambda
AWS DynamoDB
AWS API Gateway
Apache Airflow
Apache Spark
Hive
Git
Jenkins
Preferred Qualifications
Experience with EMR performance tuning and cost optimization
Experience with API design microservices or event-driven architectures
Exposure to AI/ML LLMs or AI-assisted engineering workflows
Experience with data governance data lineage or observability tools
Familiarity with real-time data streaming architectures