Senior ML Kernel Performance Engineer – Amazon – Toronto, ON
Location: Toronto, ON | Company: Amazon
At the cutting edge of AI acceleration technology, Amazon’s Annapurna Labs is looking for a Senior ML Kernel Performance Engineer to join their Acceleration Kernel Library team in Toronto, Ontario. This is a rare chance to work at the hardware-software boundary, crafting high-performance kernels for Amazon’s custom ML accelerators — Inferentia and Trainium — and directly shaping the future of deep learning and GenAI workloads.
This role sits within the broader Neuron Compiler organization, where engineers work across multiple technology layers — from frameworks and compilers to runtime and collectives. If you thrive in startup-like environments where you’re constantly solving novel problems, working cross-functionally, and making an outsized impact on a global customer base, this opportunity was built for you.
About the Role: Senior ML Kernel Performance Engineer
As a Senior ML Kernel Performance Engineer, you’ll design and implement high-performance compute kernels for ML operations using the Amazon Neuron SDK — a comprehensive toolkit that includes an ML compiler, runtime, and application framework integrating seamlessly with PyTorch and other popular ML frameworks. Your work will directly influence inference and training performance for some of the most demanding AI workloads in the world.
Beyond optimization, this role involves mentoring experienced engineers, contributing to future hardware architecture design, and collaborating directly with customers to enable their ML models. You’ll also participate in publishing cutting-edge research, making this as much an intellectual pursuit as an engineering one. Teamwork, cross-functional collaboration, and a passion for performance at scale are central to how this team operates.
Benefits and Salary
This position offers a base salary range of $150,700 to $251,700 CAD annually for the Toronto location. Amazon’s total compensation package may also include sign-on payments and restricted stock units (RSUs), with final compensation determined by experience, qualifications, and location. Benefits include comprehensive health insurance (medical, dental, vision, prescription, basic life & AD&D), a Registered Retirement Savings Plan (RRSP), Deferred Profit Sharing Plan (DPSP), paid time off, and additional resources to support health and well-being.
Job Details
🏢 Company: Amazon (Amazon Development Centre Canada ULC)
📍 Location: Toronto, ON
🆔 Job ID: 10526701
💰 Pay: $150,700 – $251,700 CAD annually
Responsibilities
Kernel engineers on this team collaborate across compiler, runtime, framework, and hardware teams to optimize machine learning workloads for a global customer base. Your day-to-day will blend low-level system optimization with customer-facing model enablement — meaning your work has both technical depth and direct business impact.
- Design and implement high-performance compute kernels for ML operations, leveraging the Neuron architecture and programming models
- Analyze and optimize kernel-level performance across multiple generations of Neuron hardware
- Profile and resolve performance bottlenecks using dedicated profiling tools and detailed analysis techniques
- Implement compiler optimizations including fusion, sharding, tiling, and scheduling strategies
- Enable and optimize customer ML models directly on AWS accelerators, providing hands-on support
- Collaborate cross-functionally to develop innovative kernel optimization techniques across teams
- Build high-impact solutions and contribute to design discussions, code reviews, and stakeholder communications
- Mentor and lead a team of experienced engineers, driving technical decisions and architecture discussions
Requirements / Skills
This role is designed for a seasoned engineer who combines deep systems knowledge with a passion for ML hardware acceleration. Amazon values candidates who can lead technically, communicate across disciplines, and move quickly in ambiguous problem spaces where there’s no established blueprint.
- 5+ years of professional software development experience outside of internships, with strong programming proficiency
- 5+ years of experience in system design or architecture, including design patterns, reliability, and scaling
- Leadership experience as a tech lead, mentor, or engineering team lead
- Expertise in accelerator architectures for ML or HPC — GPUs, CPUs, FPGAs, or custom silicon
- GPU kernel optimization experience using frameworks such as CUDA, Triton, NKI, OpenCL, SYCL, or ROCm
- Experience with NVIDIA PTX and/or AMD GPU ISA, and familiarity with LLVM/MLIR backend development
- Knowledge of ML frameworks like PyTorch or TensorFlow and their respective GPU backends
- Proficiency in parallel programming, GPU memory hierarchy optimization, and HPC library development
- A Bachelor’s degree in computer science or equivalent is preferred
How to Apply
To apply, visit the official Amazon job posting using the link below. Make sure your resume is up to date and tailored to highlight your experience with ML acceleration, kernel optimization, and systems architecture before submitting.
Share This Opportunity
Know someone who might be interested? Share this job posting and help them join Amazon in Toronto.
Job Summary & Tips for Applying
Quick Summary & What to Highlight: This Senior ML Kernel Performance Engineer role at Amazon in Toronto is perfect for candidates who excel in GPU kernel optimization, ML accelerator architecture, and compiler-level performance tuning. On your resume, emphasize any experience with CUDA, Triton, or custom silicon development, attention to low-level detail, and your ability to work across hardware and software layers in a fast-paced, high-impact environment. If you’ve previously worked in HPC, deep learning infrastructure, or ML systems engineering, make sure to highlight specific achievements and responsibilities that align with this position.
Resume & Application Tips: Before applying, tailor your resume to match the job description. Include keywords like Neuron SDK, kernel performance, and accelerator architecture that appear in the posting. Quantify your achievements where possible (e.g., “reduced kernel latency by 30% through tiling and fusion optimizations” or “led a team of 5 engineers delivering HPC library improvements”). Write a brief cover letter expressing your genuine interest in Amazon’s Annapurna Labs and why you’re excited about this opportunity in Toronto. Double-check your application for spelling errors and ensure your contact information is current.
Interview Preparation: If selected for an interview, research Amazon‘s Leadership Principles, the Neuron SDK ecosystem, and recent Inferentia/Trainium announcements beforehand. Prepare specific examples using the STAR method (Situation, Task, Action, Result) to demonstrate your system design, kernel optimization, and technical leadership skills. Common questions may include scenarios about diagnosing performance bottlenecks, leading cross-functional teams, and enabling customer workloads under constraints. Dress appropriately for a tech/engineering environment, arrive 10–15 minutes early (or log in early for virtual interviews), and bring copies of your resume. Prepare thoughtful questions about the role, team structure, and hardware roadmap. After the interview, send a thank-you email within 24 hours reiterating your interest in the position.