Enter a job title or keyword

Senior Researcher Edge AI OptimizationHardware-Aware ML


Job Location:

Edmonton - Canada

Monthly Salary: Not provided by the employer
Posted: 29 September 2026 (4 days ago)
Application Deadline: 27 December 2026
Vacancies: 1 Vacancy

Job Summary

Job description

Huawei Canada has an immediate permanent opening for a Researcher.

About the team:

The Software-Hardware System Optimization Lab focuses on research and innovation in power efficiency and performance optimization for consumer devices. By leveraging the talents and capabilities of local academia and our team we aim to build system-optimization capabilities for software and hardware across edge AI multimedia graphics mobile gaming and system software domains thereby enhancing the user experience and performance competitiveness of Huaweis consumer device products.


About the job:

  • Conduct research in hardware-aware neural network optimization (e.g. quantization-aware training mixed precision pruning distillation neural architecture search).

  • Develop novel approaches for latency/energy-aware training objectives and Pareto optimization (accuracy vs. compute vs. memory).

  • Prototype and evaluate techniques for efficient inference under device constraints (thermal limits memory bandwidth intermittent connectivity).

  • Publish and present findings internally and externally (papers workshops patents technical blogs).

  • Optimize inference pipelines across pre/post-processing scheduling operator fusion memory planning and runtime execution.

  • Collaborate on or contribute to compilers / runtimes (e.g. TVM MLIR XLA TensorRT ONNX Runtime TFLite ExecuTorch) to improve operator coverage and performance.

  • Profile and optimize models with real device traces addressing bottlenecks such as cache misses memory bandwidth kernel launch overhead and CPUNPU handoff.

  • Build and maintain hardware-aware benchmarking methodology and regression suites for edge targets (ARM CPU mobile GPU DSP NPU).

  • Create deployment recipes for heterogeneous compute (CPUGPUNPU) including partitioning strategies and fallback paths.

  • Drive optimization for on-device personalization and incremental updates when needed (e.g. small adapters efficient fine-tuning).

  • Partner with product engineering platform teams and hardware teams to translate device constraints into research targets and to transition research prototypes into production.

  • Mentor junior researchers/engineers review experimental designs and raise the quality bar for measurement rigor and reproducibility.

  • Define technical roadmap areas (e.g. next-gen quantization kernel optimization model families for edge compiler improvements).

Job requirements

About the ideal candidate:

  • PhD (or equivalent research experience) in Machine Learning Computer Science Electrical/Computer Engineering or related field.

  • Strong programming skills in Python and C/C (or equivalent systems language).Experience building AI agent / harness / skill toolchains including model evaluation orchestration and LLM-powered tooling. Hands-on experience with deep learning frameworks (e.g. PyTorch TensorFlow JAX) and deployment toolchains (e.g. ONNX TFLite TensorRT TVM MLIR-based stacks). Solid knowledge of performance profiling: latency measurement memory profiling kernel-level bottleneck analysis and experimental rigor.

  • Proven publication record at top venues (e.g. NeurIPS/ICML/ICLR MLSys ASPLOS ISCA MICRO) and/or patents in ML efficiency.

  • 2 years of relevant experience (research lab or industry) with demonstrated impact in at least one of:

    • Model compression (quantization/pruning/distillation)

    • Efficient architectures (MobileNet-like MoE efficient transformers etc.)

    • ML systems/compilers/runtime optimization

    • hardware-aware optimization for edge deployment

  • Experience optimizing for specific edge hardware:

    • ARM NEON mobile GPUs DSPs NPUs microcontrollers

    • Experience with distributed benchmarking CI for performance regression and reproducible experiment pipelines.

    • Understanding of power/thermal constraints and methodologies for measuring energy on device.

    • Experience with efficient LLM/VLM inference on edge (KV-cache optimization quantized attention speculative decoding etc.).

  • Technical Skills

    • Quantization: PTQ/QAT per-channel/per-tensor calibration smooth quant GPTQ-like methods mixed precision

    • Sparsity: structured pruning N: M sparsity hardware-friendly sparsity

    • Compiler techniques: graph rewriting operator lowering scheduling kernel autotuning

    • Runtime techniques: memory arenas tensor lifetime analysis static vs dynamic shapes batching strategies

    • Hardware fundamentals: cache hierarchy SIMD memory bandwidth accelerator programming models

Additional Information:

Huawei Canada is committed to a fair inclusive and accessible recruitment process. If you require accommodation during any stage of the hiring process please let us know and we will work with you to meet your needs.

All applications for this position are reviewed directly by our hiring team we do not use artificial intelligence tools to screen or select candidates.

All done!

Your application has been successfully submitted!

Youve already applied for this job

Thank you for your interest - weve already received your application so this new submission cant be accepted. Your previous application is on file.

If you need assistance or believe this is an error please email us at


Required Experience:

Senior IC