Data Scientist ADAS Analytics Machine Learning
San Jose, CA - USA
Job Summary
Apply vision-language models to SSR video and combine video analysis with structured telemetry to create multimodal event representations.
Build LLM classification and reasoning pipelines that triage events by severity and root cause and generate human-readable summaries of takeovers safety events and deactivations.
Design embedding pipelines and semantic search for similar-event retrieval and develop unsupervised clustering methods that discover recurring scenarios and edge-case families at fleet scale.
Build model-based ranking and scoring systems that reduce manual event triage develop active-learning loops using engineer feedback and proactively detect fleet-level anomalies.
Quantify the impact of firmware updates and configuration changes on KPIs segment driver cohorts and design A/B and quasi-experimental frameworks.
Analyze campaign effectiveness build coverage-optimization models and develop automated fleet-quality scoring.
Deploy models into PySpark and Delta Lake pipelines and the FastAPI analytics API build evaluation frameworks for foundation-model outputs and operate on the Azure data platform including ADLS Synapse and Container Apps.
- Bachelors or Masters degree in Data Science Machine Learning Statistics Computer Science or a related quantitative field. A Masters degree or PhD is beneficial but not required; demonstrated experience carries equal weight.
- 2-5 years of experience in data science or applied machine learning.
- Depth in at least one of the following: applied foundation models such as LLMs VLMs or embeddings; classical machine learning in production; or statistical experimentation.
- Strong Python skills using NumPy pandas and scikit-learn with an emphasis on clean testable code and strong SQL skills.
- Experience with ML model development including feature engineering model selection and evaluation on real data.
- Solid statistics knowledge including hypothesis testing regression and experimental design.
- Ability to communicate findings clearly to engineers and management through reports presentations and dashboards.
- Ability to turn ambiguous questions into structured analytical approaches.
Strongly Preferred
LLM or VLM application experience including prompt engineering structured outputs and evaluation.
Embedding models and vector similarity for retrieval or clustering.
PySpark for large-scale processing; candidates with strong pandas experience may ramp up.
Time-series analysis or anomaly detection.
Nice to Have
Multimodal foundation models applied to video or image data.
RAG or vector-database systems such as FAISS pgvector or Pinecone.
Spatial or geospatial clustering with DBSCAN or HDBSCAN.
Ranking or recommendation systems active learning PyTorch or TensorFlow.
Delta Lake or Parquet; FastAPI or model-serving APIs; MLOps platforms such as MLflow or Weights & Biases.
Cloud-platform experience in Azure AWS or GCP.
Vehicle telemetry or automotive-domain experience; campaign analytics or A/B testing at scale.
Data privacy including CCPA or GDPR for vehicle data and foundation models.
English required; German is an advantage.