Enter a job title or keyword

Python, PySpark India


Job Location:

Bengaluru - India

Monthly Salary: Not provided by the employer
Posted: 4 July 2026 (30+ days ago)
Application Deadline: 1 October 2026
Vacancies: 1 Vacancy

Job Summary

Total experience: 5 years

Detailed JD:
  • Build Scalable Pipelines: Design implement and maintain end-to-end batch and near real-time ETL/ELT pipelines using Python PySpark and Spark SQL.
  • Lakehouse Architecture: Implement and manage highly optimized cloud storage layers using Delta Lake or Apache Iceberg formats.
  • Orchestration & Automation: Build complex workflow schedules dependency maps and failure recovery processes using Apache Airflow.
  • Performance Tuning: Proactively debug and optimize Spark workloads by resolving data skew implementing caching strategic partitioning Z-Ordering and broadcast joins.
  • Streaming Ingestion: Construct robust real-time data ingestion pipelines using messaging systems like Apache Kafka or Cloud Pub/Sub.
  • Data Governance & Integrity: Enforce data quality validation automated schema evolution checks and structural reconciliation frameworks.


Mandatory Skills: Python PySpark