Enter a job title or keyword

Data Engineer

MSP Staffing Pty


Job Location:

Cape Town - South Africa

Monthly Salary: Not provided by the employer
Posted: 6 October 2026 (3 days ago)
Application Deadline: 3 January 2027
Vacancies: 1 Vacancy

Job Summary

Data Engineer responsible for building and maintaining data pipelines and lakehouse structures supporting analytics BI machine learning Generative AI applications and agents. The role includes Databricks delivery GenAI data enablement production-grade data engineering collaboration within cross-functional product squads and adherence to enterprise risk security and governance standards.

Education: Bachelors degree in Computer Science Engineering Information Systems or a related technical discipline. Postgraduate qualification advantageous.

Experience:

  • 6 years industry experience.
  • 6 years Senior / Lead Data Engineer experience.
  • 2 years hands-on Databricks experience.
  • 6 years enterprise data lake and lakehouse architecture.
  • 3 years Python.
  • 3 years SQL.
  • 3 years Apache Spark.
  • 3 years building and operating production-grade data platforms.
  • 5 years working in enterprise or regulated environments.

Key Requirements

  • Build and maintain data pipelines and lakehouse structures for Analytics/BI Machine Learning Generative AI applications and agents.
  • Apply enterprise data lake and lakehouse principles to ensure data is reliable well-governed secure and fit for downstream consumption.
  • Translate business and analytical requirements into production-ready data solutions.
  • Hands-on Databricks experience including Delta Lake Databricks Jobs and Workflows Unity Catalog Databricks Bundles notebooks and shared libraries.
  • Enable data consumption for GenAI use cases analytics/reporting tools and downstream operational systems.
  • Support RAG context and prompt data preparation model input/output and feedback data flows.
  • Build curated knowledge datasets structured/semi-structured data pipelines and metadata/lineage required for AI consumption.
  • Work closely with AI Engineers and Product Owners on GenAI use cases and AI Engineer development.
  • Develop production-grade pipelines using Python PySpark SQL and Apache Spark.
  • Implement automated testing and CI/CD practices for data workloads.
  • Ensure solutions are observable resilient performant and cost-efficient.
  • Contribute to data quality reliability and operational stability.
  • Collaborate with Product Owners AI/ML Engineers Analytics teams Platform and Security teams.
  • Provide engineering input into design and delivery decisions and support peer reviews and shared engineering standards.
  • Ensure compliance with enterprise security risk and governance standards.
  • Participate in incident resolution and root cause analysis.
  • Maintain appropriate documentation and runbooks.
  • Experience enabling AI ML or Generative AI use cases from a data engineering perspective.
  • Familiarity with RAG data patterns feature-style or AI-serving datasets and vector or embedding-ready data workflows.
  • Experience working in Agile product-aligned squads.
  • Exposure to cloud-native data platforms AWS or Azure.

Should you meet the requirements for this position please email your CV to . You can also contact the IT team on or visit our website at NOTE: When replying to the advert also include the reference number in the subject line. Correspondence will only be conducted with short listed candidates. Should you not hear from us within 3 days please consider your application unsuccessful.


Required Experience:

IC