Enter a job title or keyword

Senior Data Engineer, Selling Partner Agentic Interfaces Data Products

Amazon


Job Location:

Bengaluru - India

Monthly Salary: Not provided by the employer
Posted: 30 August 2026 (Yesterday)
Application Deadline: 27 November 2026
Vacancies: 1 Vacancy

Department:

Data Engineering

Job Summary

Own the architecture of a data product that makes AI-powered commerce measurable trustworthy and improvable for millions of sellers worldwide. Join SP-AI Data Products as a founding technical leader who will define how an entire organization observes governs and learns from every AI-driven interaction at scale built entirely on AWS large-scale data processing infrastructure.

Were at an inflection point. AI agents are replacing traditional seller workflows generating 7.5 million interactions annually across 50 internal product teams with thousands of Developers building on the ecosystem. Every one of these innovations creates a measurement obligation and right now theres no unified infrastructure to fulfill it. Youll change that. Youre joining a small team transforming from traditional reporting into a production data product organization and this role determines what that product becomes. Youll design and build on AWS services including Kinesis for real-time streaming ingestion EMR and Spark for large-scale distributed data processing Glue for ETL orchestration Redshift and Athena for analytical workloads S3 and Lake Formation for governed storage and Lambda and Step Functions for event-driven pipeline automation.

What youll own:

- Design and govern the canonical event schema: the unified measurement format that makes every AI interaction across every surface produce a comparable correlated record. You define what gets measured and how.
- Own end-to-end data architecture across telemetry layers (user engagement action execution domain response) ensuring cross-layer correlation through session-level tracing
- Build and operate streaming and batch ingestion pipelines on AWS with production SLAs serving real-time observability for leadership risk teams applied scientists and product teams simultaneously
- Architect tenantized data products that enable 50 teams to onboard once and receive self-serve metrics (adoption quality risk impact) without building custom pipelines
- Drive data engineering standards across the organization: schema governance naming conventions data quality operational excellence. You set the bar others build to.
- Make architectural trade-offs that balance short-term delivery against long-term scalability: tiered storage build-vs-buy decisions schema evolution across dozens of consumers
- Identify and resolve systemic architecture deficiencies proposing and leading cross-team initiatives that unblock innovation for adjacent teams
- Decompose complex ambiguous problems into parallel workstreams executable by you and others then reassemble them into cohesive solutions
- Elevate the engineering team through mentorship and technical leadership. Your presence makes the team stronger but the team doesnt require your presence to succeed.

Why youll love this role:

- Founding-team impact: Youre defining the architecture an entire organization builds on not inheriting legacy systems
- Breadth of influence: Your schema decisions quality standards and architectural patterns are consumed by risk teams scientists product managers and leadership across the business
- Technical depth at scale: Streaming infrastructure cross-surface correlation schema governance for 50 consumers production SLAs on AWS. Hard consequential engineering.
- Career-defining scope: Cross-cutting schema design and multi-team domain onboarding at this scale is the kind of work that shapes what comes next

If youve built large-scale data products on AWS that other teams depend on thrive in ambiguity and want to define how an organization measures AI-driven commerce at scale wed love to talk.

Key job responsibilities
- Own the design implementation and evolution of large-scale data architecture on AWS (Kinesis EMR Spark Glue Redshift Athena S3 Lake Formation) providing system-wide technical guidance and ensuring all data products meet production-grade reliability scalability and security standards
- Write exemplary production-quality code in Java Scala and Python to build streaming and batch data pipelines that process terabytes of telemetry data daily across distributed systems with proper testing lineage documentation and quality controls. SQL proficiency is expected as a baseline.
- Design and maintain the canonical event schema and standardized reusable data products following data mesh governance principles that eliminate metric inconsistencies and enable self-service analytics for 50 domain teams
- Build and operate ETL and ELT pipelines using AWS-native services (Glue EMR Step Functions Lambda) and JVM-based frameworks (Spark on Scala/Java) that transform raw telemetry into trusted AI-ready datasets with defined SLAs for freshness and completeness
- Architect real-time observability infrastructure and analytical workloads (Kinesis Redshift Athena QuickSight) that provide actionable insights on user engagement operational health risk detection quality monitoring and compliance tracking
- Collaborate with and influence peer Software Development Engineers Applied Scientists Product Managers and Business Intelligence Engineers to define requirements validate data quality and align data product design with business and technical strategy
- Define and enforce data governance standards across the organization including schema governance Fine-Grained Access Control data classification naming conventions and lineage tracking across all data products
- Identify systemic data quality issues and architecture deficiencies drive root-cause resolution and automate manual processes to improve operational efficiency and eliminate recurring failures
- Document data products architectural decisions and design patterns clearly to ensure ease of use extensibility and maintainability by team members and downstream consumers across the organization
- Lead cross-team initiatives to enable GenAI capabilities through Amazon Quick Suite defining how AI surfaces consume governed data while maintaining robust security access controls and auditability at scale

A day in the life
Your primary focus is owning the data architecture that powers AI-driven commerce for millions of sellers. Most of your day is spent writing Java Scala and Python code: building Spark pipelines on EMR designing streaming ingestion on Kinesis and optimizing how terabytes of interaction data flow through AWS infrastructure into governed queryable products.

Youll be based in Bengaluru working alongside a team split between Bengaluru and Seattle. Your mornings give you uninterrupted build time while Seattle is offline. You use this window for deep architectural work: writing design documents pushing complex pipeline code and leaving detailed code review feedback so your Seattle teammates wake up unblocked. During overlap hours you join focused design reviews and cross-team discussions where your job is to bring clarity shape schema decisions and drive alignment across engineering teams consuming your data products.

Beyond the core technical work youre the person domain teams come to when they need guidance on how to emit telemetry that conforms to the canonical schema. Youre investigating why a streaming pipeline breached its SLA and fixing the root cause before it recurs. Youre mentoring engineers through pull requests that teach long-term maintainability not just correctness. Some days bring unexpected production issues that need fast diagnosis. Other days youre heads-down on a multi-week design for the next generation of tenantized data products that will serve 50 teams.

The Bengaluru-Seattle structure means you operate with high autonomy. Decisions dont wait for handoffs across timezones. You own outcomes end-to-end and the team trusts your judgment to move fast and get it right.

About the team
SP-AI Data Products is a small high-autonomy team that builds the data infrastructure powering AI-driven commerce for millions of Amazon Selling Partners worldwide. Our mission is straightforward: make every AI agent interaction measurable governable and improvable through trusted self-service data products. We exist so that product teams can ship with confidence risk teams can detect abuse in real time scientists can train better models and leadership can see exactly how the business is performing without waiting for someone to compile a report.

Were a team of software engineers data engineers and business intelligence engineers split between Bengaluru and Seattle. We operate like a startup inside Amazon: small enough that every persons work is visible and consequential but connected to a business growing at 7x year-over-year that serves 50 domain teams thousands third-party developers and millions of sellers. Our customers are internal: the engineering teams building AI agents the risk and trust teams protecting the ecosystem the product managers tracking adoption and the scientists improving model quality. We build for all of them simultaneously through governed tenantized data products rather than one-off reports.

Our culture is built on ownership and craft. We write production-grade code in Java Scala and Python. We design schemas that dozens of teams consume. We hold each other to high engineering standards through rigorous code and design reviews and we invest in mentorship because raising the bar across the team matters more than any individual contribution. Were transparent about what we are: a lean team executing an ambitious transformation from reactive reporting into a production data product organization built entirely on AWS. If you want to work somewhere your architectural decisions directly shape how an organization operates where youll have real autonomy to drive outcomes and where the problems are genuinely hard and unsolved this is the team.

- 7 years of data engineering experience
- 5 years of development/programming/scripting language (Python/Java/Bash/Perl) experience
- Experience with AWS services including S3 Redshift Sagemaker EMR Kinesis Lambda and EC2
- Experience with data modeling warehousing and building ETL pipelines
- Experience leading engineering teams as a mentor or tech lead or experience leading the architecture and design (architecture design patterns reliability and scaling) of new and current systems

- Experience in any Bigdata architecture or experience managing full application stacks from the OS up through custom applications and experience with automation and any version control tools
- Experience in Kafka or experience in any Bigdata architecture and experience in Redshift
- Experience in data warehouse technical architectures data modeling infrastructure components ETL/ ELT and reporting/analytic tools and environments data structures and hands-on SQL coding

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process including support for the interview or onboarding process please visit for more information. If the country/region youre applying in isnt listed please contact your Recruiting Partner.


Required Experience:

Senior IC


About Company

Company Logo

Free shipping on millions of items. Get the best of Shopping and Entertainment with Prime. Enjoy low prices and great deals on the largest selection of everyday essentials and other products, including fashion, home, beauty, electronics, Alexa Devices, sporting goods, toys, automotive ... View more

View Profile View Profile