PhD Research Scientist Intern Edge AI
Job Summary
At Canva our mission is to empower the world to design. Were building AI that feels magical and lands real impact for millions of people helping anyone create with confidence. Were looking for a research intern who is excited by efficient ML and edge deployment to help us bring video-capable vision-language models onto the devices in peoples pockets.
About the team
Were the Video Storytelling team working on the models and systems behind Canvas video AI experiences. We partner closely with our Edge AI group who are building Canvas on-device inference capability to explore whats possible when AI runs directly on users own hardware. We already own several of the server-side capabilities this work builds on so youll be joining a team with a strong command of the data models and pipelines behind the problem.
About the role
This is a 14-week research internship focuses on one clear question: can a video-capable vision-language model be optimised to run efficiently on high-traffic consumer phones while retaining enough capability to serve a real product use case
The use case is intelligent captioning where the models visual understanding of a video drives context-aware intelligently placed captions. Youll own the complete arc from model selection through optimisation deployment and measurement delivering a working on-device prototype plus a benchmarked map of what current consumer hardware can and cant do. Youll inherit mature data and pipelines from our work so you can benchmark directly against a strong reference rather than building from scratch. Youll be supported by supervisors with deep on-device AI backgrounds weekly 1:1s and collaborators across teams.
What youll do
Survey candidate video-capable VLMs (e.g. Gemma Qwen-VL SmolVLM MiniCPM-V) and determine the best starting point
Apply model optimization techniques and architecture improvements to specialize vision-language models for on-device deployment including quantization pruning distillation hardware-specific compilation and task-specific fine-tuning for caption placement.
Deploy the model on real high-traffic mobile hardware through our on-device inference library iterating the optimisation-deployment loop against real on-device measurements.
Run comparative evaluation against at least one alternative optimisation path and human evaluation against our server-side captions quality bar.
Document your findings clearly enough that the team can act on them mapping which workloads are viable on-device today and which arent yet and why.
Compile your output into a patent filing and a paper publication.
Youre likely a match if you have
Strong Python and hands-on PyTorch experience including training and fine-tuning vision-language models.
A solid understanding of modern vision-language and multimodal architectures with the ability to pick up a recent paper and reproduce it.
Experience with optimisation methods like quantisation pruning or distillation and a clear sense of what each costs you in accuracy.
Experience deploying models on-device or at the edge with runtimes like Core ML LiteRT/TFLite ONNX Runtime or ExecuTorch working within real memory and latency budgets.
Experience running your own research project end to end: making a plan measuring carefully and iterating on what you find.
Current enrolment in a PhD in ML CS or a related field with first-author papers at venues like CVPR NeurIPS ICCV/ECCV ICLR or ICML.
Nice to have
Experience with video understanding models ideally the token-efficient kind.
Publications or open-source contributions in efficient ML multimodal models or edge AI.
Experience writing custom kernels for inference optimisation.
Experience deploying models across different on-device hardware accelerators (e.g. Apple Neural Engine DSPs).
Experience working across research and product teams in industry or on a previous internship.
Additional Information :
Other stuff to know
We make hiring decisions based on your experience skills and passion as well as how you can enhance Canva and our culture. When you apply please tell us the pronouns you use and any reasonable adjustments you may need during the interview process.
We celebrate all types of skills and backgrounds at Canva so even if you dont feel like your skills quite match whats listed above - we still want to hear from you!
Please note that interviews are conducted virtually.
Remote Work :
No
Employment Type :
Intern
About Company
We're a global online visual communications platform on a mission to empower the world to design. Featuring a simple drag-and-drop user interface and a vast range of templates ranging from presentations, documents, websites, social media graphics, posters, apparel to videos, plus a hu ... View more