Senior Applied Data Scientist | NDA
Job Summary
GT was founded in 2019 by a former Apple Nest and Google executive.GTs mission is to connect the worlds best talent with product careers offered by high-growth companies in the UK USA Canada Germany and the Netherlands.
On behalf of our client GT is looking for a Senior Applied Data Scientist interested in developing and testing new ML embedding and LLM-based approaches to solve complex data matching problems at scale.
Our client is a leading global management consultancy known for tackling some of the worlds most complex business challenges. With a focus on strategy transformation and performance improvement the firm partners with major organizations across industries to drive lasting impact.
We are looking for a Senior Applied Data Scientist to improve how entity resolution is performed at scale.
You will develop and test new ML embedding and LLM-based approaches for matching complex business records across multiple data sources.
The work is centered on model quality experimentation and evaluation; engineering partners will help productionize successful approaches.
A key part of the role is exploring how newer foundation-model techniques can improve matching quality while remaining practical and scalable for very large datasets.
Develop better ways to match company records
Build new ML embedding and LLM-based approaches for matching entities
Improve how the system handles messy data including name variations aliases domains websites firmographic attributes multilingual records and data hierarchies.
Develop scoring and ranking approaches to distinguish accurate matches from duplicates similar-looking records and unrelated entities.
Evaluate and implement AI and machine learning techniques to improve matching quality while considering accuracy scalability and cost.
Design approaches that can operate efficiently at scale taking model usage and computational cost into consideration.
Improve evaluation experimentation and match quality
Define and improve methods for evaluating match quality including precision recall false positives false negatives confidence coverage and manual review effort.
Assist in building trusted benchmark sets that allow us to compare new models against the current matching engine before production rollout.
Explore LLM-assisted review and validation to assess matching performance and benchmark more scalable approaches.
Turn ambiguous matching problems into clear hypotheses experiments metrics and recommendations.
Partner with engineering to bring successful ideas into production
Work closely with data engineering and software engineering teams to turn promising prototypes into production-ready matching logic.
Provide engineering partners with clear model specifications evaluation results expected behavior edge cases and rollout requirements.
Help determine the most appropriate matching techniques based on data characteristics confidence levels and cost considerations.
Continuously evaluate matching performance investigate regressions and recommend improvements to models and matching logic.
Clearly communicate technical tradeoffs related to matching performance scalability cost latency explainability and operational considerations.
58 years of relevant experience in Data Science Applied Data Science Applied Machine Learning or a similar role.
Strong applied ML fundamentals with hands-on experience building and evaluating models on real data.
Excellent Python and SQL skills.
Practical experience with embeddings semantic similarity LLMs or related AI techniques.
Hands-on experience training supervised and unsupervised models including classification and NLP tasks.
Working knowledge of neural network and transformer architectures.
Proficiency with common ML frameworks such as TensorFlow PyTorch and PyCaret.
Experience retraining a taxonomy classifier or maintaining classification models in production.
Experimental judgment: able to define baselines metrics test sets and error analysis that show whether quality improved.
Ability to explain model behavior tradeoffs and edge cases clearly to engineering and business partners.
Experience with entity resolution record linkage deduplication or similar matching problems.
Experience with ranking similarity scoring retrieval clustering or candidate generation.
Experience applying LLMs or embeddings to business problems where cost and scale matter.
Exposure to large-scale data platforms such as Spark Snowflake Databricks or BigQuery.
Familiarity with company domain website firmographic or other business-entity data.
GT interview with Recruiter
Technical interview
Final interview
Required Experience:
Senior IC
About Company
GT provides high-growth product companies around the world with offshore product teams from Eastern Europe, an end-to-end product development studio, software development, and data science services.