Data Engineer
Job Summary
At Preply were all about creating life-changing learning experiences. We help people discover the magic of the perfect tutor craft a personalised learning journey and stay motivated to keep growing. Our approach is human-led tech-enabled - and its creating real impact.
Weve just reached unicorn status with a $150M Series D accelerating our vision to transform education through human-led AI-enhanced learning. Today 100000 tutors teach 90 languages to learners in 180 countries - and were only getting started. As a category-defining company were shaping what the future of learning looks like at global scale.
Every Preply lesson sparks change fuels ambition and drives progress that matters. Joining Preply means helping define the future of education at global scale and building something that truly matters for millions of people every day.
At Preply the Data Ingestion and Enrichment team provides a single trusted and scalable data foundation. The team ensures that all analytics machine learning and product features are built on unified governed and production-grade data assets in Preplys Lake House including the extraction normalization and generation of structured data from Preplys unstructured assets forming a durable data moat for AI-driven products.
As a Data Engineer in the Data Ingestion and Enrichment team you will build and contribute to the data layer that powers both Preplys analytics machine learning and product. You will work closely with ML Platform Applied/Data Scientists Analytics Engineering and Product squads to ensure that features datasets and pipelines are production-ready observable and reusable within the team.
Contribute to trusted ingestion & enrichment foundations (Data Lake and Data as a Product):
Build and maintain components of Preplys data lake. Ensure every dataset has clear ownership purpose schemas and quality expectations from first ingestion through downstream consumption by analytics product and ML teams. Treat trust correctness and predictability as first-class features of the platform.
Develop end-to-end ingestion pipelines (batch & streaming):
Build and operate reliable batch and streaming ingestion pipelines that support both real-time and analytical use cases. Contribute to defining clear raw standardized consumption layers with explicit responsibilities lineage and retention strategies. Balance performance cost and reliability as the platform scales.
Data quality contracts & early validation:
Implement data contracts between producers and consumers covering schema freshness volume and quality guarantees. Embed validation anomaly detection and quality checks early in the ingestion lifecycle to catch issues before they propagate. Apply standardized quality metrics.
Enrichment modeling & lifecycle management:
Build enrichment logic that joins standardizes and contextualizes data across domains using shared definitions and reusable patterns. Support historical tracking point-in-time correctness and dataset versioning so downstream users can confidently analyze changes and impacts over time.
Observability reliability & operational excellence:
Instrument ingestion pipelines with strong observability: freshness latency data quality and cost metrics. Contribute to SLOs alerting and incident response playbooks so data failures are visible diagnosable and recoverable. Help move the platform from reactive firefighting to proactive reliability management.
Governance & compliance by design:
Apply consistent access control classification and privacy protections at ingestion time. Ensure sensitive data is properly masked minimized or anonymized by default and that all data flows you own are auditable and traceable.
Enable self-service & standardization:
Contribute to standardized ingestion templates shared libraries and platform tooling that enable teams to onboard new data sources independently. Improve discoverability documentation and metadata so datasets you own are easy to find and trust without relying on tribal knowledge.
Cross-team collaboration & ownership:
Work closely with Product Backend Analytics and ML partners to align on ingestion requirements and trade-offs. Build strong working relationships across teams. Mentor junior team members and actively contribute to a culture of shared data quality standards and data contracts.
Hands-on experience building components of large high-scale applications (e.g. data pipelines well-structured APIs efficient algorithms).
Solid experience working in platform or data engineering teams (or equivalent) with the ability to deliver within a multi-stakeholder environment.
Familiarity with cloud platforms (AWS/GCP or equivalent) and modern DevOps practices.
Hands-on experience designing and implementing real-time and batch data processing pipelines using modern frameworks like Spark Flink Spark Streaming Kafka Debezium etc.
Experience with orchestration tools such as Airflow dbt or similar.
Exceptional problem-solving skills paired with a proactive innovative mindset focused on continuous improvement.
Strong communication and cross-functional collaboration skills (English level B2)
An open collaborative dynamic and diverse culture;
A generous monthly allowance for lessons on Learning & Development budget and time off for your self-development.
A competitive financial package with equity leave allowance and health insurance;
Access to free mental health support platforms;
The opportunity to unlock the potential of learners and tutors through language learning and teaching in 175 countries (and counting!).
Care to change the world - We are passionate about our work and care deeply about its impact to be life changing.
We do it for learners - For both Preply and tutors learners are why we do what we do. Every day we focus on empowering tutors to deliver an exceptional learning experience.
Keep perfecting - To create an outstanding customer experience we focus on simplicity smoothness and enjoyment continually perfecting it as every detail matters.
Now is the time - In a fast-paced world it matters how quickly we act. Now is the time to make great things happen.
Disciplined execution - What makes us disciplined is the excellence in our execution. We set clear goals focus on what matters and utilize our resources efficiently.
Dive deep - We leverage business acumen and curiosity to investigate disparities between numbers and stories unlocking meaningful insights to guide our decisions.
Growth mindset - We proactively seek growth opportunities and believe todays best performance becomes tomorrows starting point. We humbly embrace feedback and learn from setbacks.
Raise the bar - We raise our performance standards continuously alongside each new hire and promotion. We build diverse and high-performing teams that can make a real difference.
Challenge disagree and commit - We value open and candid communication even when we dont fully agree. We speak our minds challenge when necessary and fully commit to decisions once made.
One Preply - We prioritize collaboration inclusion and the success of our team over personal ambitions. Together we support and celebrate each others progress.
Required Experience:
IC