Google Data Platform Engineer
Los Gatos, CA - USA
Job Summary
4-8 years data engineering including 2 years hands-on GCP. Reports to the Senior Platform Engineer.
A build role against a defined architecture. Implements assigned pipelines and models to the standard set by the architect and senior engineer. Streaming is the default on this platform not an occasional requirement.
-
Build streaming Dataflow pipelines in Apache Beam consuming Pub/Sub events into the BigQuery bronze layer.
-
Implement deduplication idempotent writes event ordering and late-arriving event handling the logic most likely to fail silently if done carelessly.
-
Implement schemas as data contracts and handle schema evolution without dropping or corrupting events.
-
Implement DLQ routing message archival and the replay path and test recovery under realistic failure rather than happy-path only.
-
Build Dataform models across conformed and mart layers with meaningful assertions plus business-friendly table and column documentation as part of the build.
-
Apply BigQuery performance and cost practices in code: partitioning clustering incremental materializations.
-
Build reconciliation checks against the system of record and produce sign-off evidence.
-
Register datasets in Dataplex and apply policy tags and row-level security to the required granularity.
-
Build one-time historical migration loads from files and database extracts reconciled against the streaming path at cutover.
-
Contribute Terraform modules and CI/CD; write tests including replay and duplicate-event scenarios.
-
Emit structured logs and metrics from every pipeline so the platforms operations layer can monitor it; write runbooks; support UAT cutover and hypercare.
-
Hands-on streaming experience Pub/Sub and Dataflow or Kafka / Flink / Kinesis with real exposure to deduplication ordering and replay.
-
Strong Python and advanced SQL: window functions CTEs incremental merge patterns query tuning.
-
Apache Beam or demonstrable ability to ramp quickly from another streaming framework.
-
Hands-on BigQuery: partitioning clustering cost-aware query design.
-
Dataform or dbt including tests or assertions and dependency management.
-
Working knowledge of dimensional modeling.
-
Git workflow and CI/CD; Terraform or willingness to ramp quickly.
-
Exposure to a major SaaS platform as a data source and its change-event mechanisms.
-
Comfort building to an architecture someone else defined raising concerns through the right channel rather than deviating quietly.
-
GCP Professional Data Engineer certification; Dataplex and DLP; Analytics Hub or Looker familiarity.