Data Engineer
Job Summary
Keep millions of daily transactions clean trustworthy and moving without turning every upstream change into a fire drill.
Paris France Permanent On-site
- Building and owning batch and streaming pipelines that feed the companys core data warehouse
- Ingesting data from transactional databases internal services APIs and event streams
- Designing schemas and data models that remain consistent as the business adds new products and markets
- Developing and maintaining reusable transformation models in dbt
- Setting up automated checks for freshness completeness duplicates schema changes and business consistency
- Monitoring production pipelines and investigating failures across sources orchestration transformations and warehouse workloads
- Managing retries late-arriving data backfills and historical reprocessing
- Optimizing pipeline performance query execution storage and compute costs as data volumes increase
- Planning schema migrations without disrupting downstream users
- Partnering directly with data scientists and analysts to translate their requirements into reliable datasets
- Documenting lineage dependencies transformation logic and ownership so critical data can be understood and trusted
- Contributing to code reviews automated testing CI/CD and data engineering standards
- Streaming ingestion with Kafka alongside batch orchestration with Airflow with clear decisions about when each approach is appropriate
- A Snowflake or BigQuery warehouse under real query and transformation load from multiple internal teams
- A growing dbt transformation layer that needs to remain tested documented and maintainable
- Incremental processing across large datasets where full refreshes are not a realistic option
- Backfills and schema migrations on tables that cannot simply be rerun from scratch
- Handling partial failures upstream changes delayed events and dependencies between critical pipelines
- Improving data quality without producing large volumes of low-value alerts
- Balancing reliability and performance with infrastructure and warehouse costs
- 4 years of experience building and operating production data pipelines
- Strong SQL and Python skills
- Practical experience with Airflow dbt Kafka Snowflake BigQuery or similar technologies
- Good knowledge of data modelling and incremental processing
- Experience with automated testing monitoring and production troubleshooting
- An understanding of retries partial failures late-arriving data backfills and pipeline dependencies
- The ability to investigate issues across the full data flow
- Experience using Git code reviews and CI/CD
- Confidence working directly with analysts data scientists and engineering teams
- Ownership of data quality not just the movement of data between systems
You do not need to have worked with every tool in the stack but you should understand the engineering principles behind reliable and maintainable data systems.
A fast-growing e-commerce platform operating across several European markets with several hundred employees and millions of transactions processed every month.
The data platform supports reporting product analytics finance operations and data science. The team is now scaling its pipelines and models to support higher volumes new markets and a growing number of internal users.
Health insurance meal vouchers and a remote-friendly policy.
Languages: Native or bilingual French and professional English.
A search run by The French Sourcer recruitment built for technical teams.