Enter a job title or keyword

Senior Software Engineer, Perception (Robotics)

Grab


Job Location:

Shanghai - China

Monthly Salary: Not provided by the employer
Posted: 21 August 2026 (11 hours ago)
Application Deadline: 18 November 2026
Vacancies: 1 Vacancy

Job Summary

About the Team:

The Robotics Technology team is a core part of Grabs long-term vision to build urban embodied AI. Our engineers take full ownership of the product lifecycle: designing and manufacturing hardware in-house developing control and machine-learning systems and rigorously testing in real-world conditions and production fleet operations. We are executing an ambitious growth plan to expand our robotics fleet across cities over the coming years and we are focused on delivering highly productive safe and efficient robot delivery services that help address current delivery labour shortages.

Based in Singapore and China we offer opportunities to work on the latest autonomy deploy solutions in complex environments and directly influence the future of last-mile logistics. If youre excited by tangible impact large-scale systems and cross-functional engineering youll find meaningful challenges and rapid career growth here.

Get to know the role

As a Senior Perception & Prediction Engineer you will build and ship systems that turn raw multi-sensor data into a reliable real-time and predictive representation of the world. You will work across a robust modular perception stack and modern learning-based approaches contributing through model development evaluation debugging integration and on-robot validation.

On top of a shipped multi-sensor detection baseline you will deepen three connected capabilities: multi-object tracking generic / open-set world understanding and motion prediction. You will also help evaluate and productionise end-to-end temporal perception-prediction models and VLA / embodied foundation models for open-vocabulary understanding long-tail reasoning and task-conditioned robot intelligence.

We are pragmatic about new research: a model earns its place through measurable closed-loop value reliable grounding real-time performance and safety. Classical geometry filtering and modular components remain important as interpretable baselines safety fallbacks and guardrails. This role offers the opportunity to take promising research from prototype to fleet data embedded deployment and real-world robot behaviour.

You will report to the Senior Principal Perception & Prediction Engineer and work onsite at a Grab office.

The critical tasks you will perform

  • You will develop and improve multi-object tracking data association Bayesian state estimation (Kalman / EKF / UKF and motion models) track lifecycle and ID stability through detector gaps occlusions and crowded scenes and evolve it toward learning-based tracking including learned association joint detection-and-tracking and query / transformer-based trackers.
  • You will build and harden generic / open-set world understanding class-agnostic obstacle detection occupancy grids / occupancy networks / BEV occupancy clustering and learned generic-object branches so the robot safely reacts to rare or unknown objects.
  • You will build motion prediction for pedestrians cyclists vehicles and other agents including multi-modal trajectory forecasting interaction-aware prediction occupancy flow uncertainty estimation and well-defined interfaces into behaviour and planning.
  • You will develop and evaluate end-to-end temporal perception-prediction approaches such as joint detection-tracking-forecasting streaming BEV representations agent / trajectory queries and learned world models while maintaining strong modular baselines and production fallbacks.
  • You will explore VLM / VLA and embodied foundation models for open-vocabulary perception semantic scene reasoning task-conditioned understanding long-tail discovery auto-labeling and teacher-student supervision. You will test grounding hallucination temporal consistency latency and robustness then distil quantise or adapt useful capabilities for on-robot deployment rather than stopping at demos.
  • You will strengthen multi-sensor and temporal integration across camera LiDAR radar and IMU ensuring geometry timing identities predictions and semantic context remain consistent over time.
  • You will define quality gates across tracking (MOTA / MOTP-style ID-switch fragmentation) prediction (ADE / FDE miss rate NLL / calibration) generic-object / occupancy / occupancy-flow recall closed-loop safety outcomes and embedded latency budgets and drive down regressions with fleet-log root-cause analysis.
  • You will partner with Detection Prediction Planning Data & Infra and Integration teams to move classical learning-based and foundation-model capabilities from research into reproducible training replay simulation NVIDIA Orin / TensorRT deployment and production-like robot testing.

Qualifications :

Skills you need

What Essential Skills You Will Need

  • At least 3 years hands-on experience shipping or operating perception and / or prediction systems for robotics AV or ADAS with strong autonomous-driving experience and a record of improving real-world performance.
  • Rich experience across classical / geometric and learning-based approaches with sound judgement on when to use modular methods end-to-end models or hybrid designs and how to retain interpretable safety fallbacks.
  • Proven depth in multi-object tracking data association Bayesian filtering track lifecycle multi-sensor / temporal association ID stability and low-latency evaluation; experience with learned or query-based tracking is valuable.
  • Hands-on experience with generic / open-set perception class-agnostic obstacle detection occupancy representations LiDAR clustering anomaly / unknown-object handling or fusion with learned detections.
  • Hands-on depth in at least one modern temporal area: motion forecasting occupancy flow streaming / temporal BEV learned world models or joint detection-tracking-prediction. You should understand multi-modal uncertainty and the interface between perception prediction and planning.
  • Fundamentals in geometry and multi-sensor systems including practical calibration know-how and the ability to diagnose alignment or time-synchronisation issues that impact model quality.
  • Proficiency in C and Python with the ability to write production-quality code build training / evaluation tools and reason about real-time embedded performance.
  • Experience navigating noisy production-like environments through reproducible experiments clear metrics log-driven root-cause analysis simulation / replay and validation-minded execution.
  • Self-motivated and dependable with genuine enthusiasm for autonomous driving robotics and embodied AI and the judgement to separate robust progress from research hype.
  • Demonstrated proficiency in leveraging AI tools and systems to accelerate engineering and research without compromising quality security or reproducibility.

Good to have

  • End-to-end perception and prediction: joint detection-tracking-forecasting occupancy / occupancy-flow prediction streaming BEV agent / trajectory queries or world models.
  • Motion prediction for autonomous systems: multi-modal trajectories interaction-aware forecasting uncertainty calibration planning interfaces closed-loop simulation or behaviour evaluation.
  • Experience with VLM / VLA / embodied foundation models including vision-language grounding open-vocabulary perception task or instruction conditioning semantic scene reasoning long-tail mining or action-relevant representations.
  • Foundation-model adaptation and deployment: supervised fine-tuning / LoRA preference or imitation data distillation quantisation teacher-student pipelines edge inference and systematic evaluation of hallucination and grounding.
  • 3D / BEV detection and sensor fusion (for example BEVFusion / LSS) temporal models 4D data and multi-camera / LiDAR / radar integration.
  • Experience with common robotics / perception libraries and platforms such as OpenCV PCL Ceres ROS 2 NVIDIA Orin TensorRT or DeepStream plus strong visualisation and debugging practices.
  • Familiarity with the model lifecycle: dataset curation auto-labeling temporal annotation active learning regression detection inference optimisation and fleet-data feedback loops.

Additional Information :

Life at Grab

We care about your well-being at Grab here are some of the global benefits we offer:

  • We have your back with Term Life Insurance and comprehensive Medical Insurance.
  • With GrabFlex create a benefits package that suits your needs and aspirations.
  • Celebrate moments that matter in life with loved ones through Parental and Birthday leave and give back to your communities through Love-all-Serve-all (LASA) volunteering leave
  • We have a confidential Grabber Assistance Programme to guide and uplift you and your loved ones through lifes challenges.
  • Balancing personal commitments and lifes demands are made easier with our FlexWork arrangements such as differentiated hours

What We Stand For at Grab

We are committed to building an inclusive and equitable workplace that enables diverse Grabbers to grow and perform at their best. As an equal opportunity employer we consider all candidates fairly and equally regardless of nationality ethnicity religion age gender identity sexual orientation family commitments physical and mental impairments or disabilities and other attributes that make them unique.


Remote Work :

No


Employment Type :

Full-time


About Company

Company Logo

About Grab and Our WorkplaceGrab is Southeast Asia's leading superapp. From getting your favourite meals delivered to helping you manage your finances and getting around town hassle-free, we've got your back with everything. In Grab, purpose gives us joy and habits build excellence, w ... View more

View Profile View Profile