JobFlexy

Sr. SDE, Edge AI ML Platform – Amazon – Vancouver, BC

Location: Vancouver, BC | Company: Amazon

Amazon’s Edge AI and Science team is hiring a Senior Software Development Engineer to lead the architecture and delivery of a cutting-edge ML platform in Vancouver, BC. This is a high-impact role within the Edge AI ML Platform and Infrastructure team — the group responsible for building the systems that let Amazon teams train, optimize, evaluate, and deploy generative AI models on devices and in the cloud.

Sponsored Links

If you’re energized by solving hard problems at the intersection of distributed systems, GPU performance, and machine learning infrastructure, this role puts you at the centre of it. You’ll work alongside applied scientists, ML engineers, GPU kernel engineers, compiler teams, and hardware teams to deliver platforms that support models with hundreds of billions of parameters.

About the Role: Senior Software Development Engineer, Edge AI ML Platform

This is a technical leadership role that blends hands-on engineering with broader architectural ownership. You’ll be responsible for designing and delivering platform capabilities across model ingestion, distributed training, compression, evaluation, and deployment. The team’s goal is to turn what today requires expert manual coordination into a repeatable, self-service workflow — and you’ll be one of the principal engineers making that happen.

Beyond the code, you’ll define stable APIs and architecture boundaries, lead cross-team programs, mentor engineers, and help build a strong engineering culture in Vancouver. You’ll use performance, reliability, and developer productivity data to drive platform investments and ensure the team is solving problems at their root rather than patching symptoms.

Sponsored Links

Benefits and Salary

Amazon offers a competitive compensation package for this role. The base salary range for Vancouver, BC is $150,700 to $251,700 CAD annually. Final compensation is determined based on experience, qualifications, and location. In addition to base pay, the total compensation package may include sign-on payments and restricted stock units (RSUs). Amazon also provides comprehensive benefits including health insurance (medical, dental, vision, prescription, basic life and AD&D), a Registered Retirement Savings Plan (RRSP), a Deferred Profit Sharing Plan (DPSP), paid time off, and additional resources to support health and well-being.

Job Details

📌 Job Type: Full-Time

🏢 Company: Amazon Development Centre Canada ULC

📍 Location: Vancouver, BC

🆔 Requisition ID: 10492307

💰 Pay: $150,700 – $251,700 CAD annually

Responsibilities

In this role, you’ll be working across the full ML platform lifecycle — from model onboarding through to production deployment. Your day will move between reviewing architecture designs, investigating distributed training failures, profiling GPU workloads, and driving cross-team delivery. These responsibilities span both deep technical execution and strategic platform leadership.

  • Lead the design and delivery of distributed ML platform services and libraries across model ingestion, optimization, training, evaluation, packaging, and deployment
  • Define stable APIs and architecture boundaries that allow scientists to contribute algorithms without coupling research code to infrastructure or deployment implementations
  • Design distributed training capabilities across data, tensor, pipeline, and model parallelism for large language and multimodal models
  • Scale workflows on multi-node GPU clusters while improving training throughput, GPU utilization, memory efficiency, communication performance, and failure recovery
  • Develop infrastructure connecting distributed training with distillation, quantization, pruning, and other model optimization techniques
  • Build evaluation and artifact workflows that measure model quality and system performance, carrying validated models through deployment on target hardware
  • Establish automated CI/CD, regression testing, observability, and release mechanisms for GPU-intensive ML workloads
  • Profile and optimize end-to-end system performance with applied scientists and GPU kernel engineers, translating bottlenecks into durable platform improvements
  • Partner with model, compiler, runtime, hardware, security, and infrastructure teams to manage technical dependencies and deliver multi-team programs
  • Mentor engineers, improve code and design review practices, and contribute to recruiting and team development in Vancouver

Requirements / Skills

Amazon is looking for a senior engineer who brings both deep technical expertise in distributed systems and ML infrastructure and the leadership instincts to guide complex, multi-team initiatives. You should be comfortable moving between high-level architecture decisions and low-level implementation details, and you bring a track record of delivering reliable, scalable systems in demanding environments.

  • 5+ years of professional software development experience (non-internship), with strong proficiency in at least one programming language
  • 5+ years of experience leading design or architecture of new and existing systems, including design patterns, reliability, and scaling
  • Experience as a tech lead or engineering team lead, with demonstrated ability to mentor engineers and drive technical direction
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field
  • Experience designing or building distributed systems or high-performance computing systems
  • Experience with distributed ML training frameworks such as PyTorch, TensorFlow, JAX, NeMo, or Megatron (preferred)
  • Familiarity with CUDA kernels, ML/low-level kernels, or profiling and debugging large-scale systems (preferred)
  • Experience with containers, Kubernetes, AWS infrastructure, CI/CD, and production operations (preferred)
  • Knowledge of model compression, quantization, knowledge distillation, or edge deployment techniques (preferred)

How to Apply

To apply, visit the official Amazon job posting using the link below. Make sure your resume is up to date and reflects your experience with distributed systems, ML infrastructure, and technical leadership before submitting.

Share This Opportunity

Know someone who might be interested? Share this job posting and help them join Amazon in Vancouver.

Job Summary & Tips for Applying

AI-generated summary and tips to help you highlight your strengths effectively.

Quick Summary & What to Highlight: This Senior Software Development Engineer role at Amazon in Vancouver is perfect for candidates who excel in distributed systems design, ML infrastructure engineering, and technical leadership. On your resume, emphasize any experience with large-scale GPU training platforms, model optimization pipelines, and cross-functional team delivery. If you’ve previously worked in ML infrastructure, cloud platform engineering, or high-performance computing, make sure to highlight specific achievements and responsibilities that align with this position.

Resume & Application Tips: Before applying, tailor your resume to match the job description. Include keywords like distributed training, model optimization, and GPU performance that appear in the posting. Quantify your achievements where possible (e.g., “reduced model training time by 30% through pipeline parallelism” or “led architecture for a platform serving 10+ model teams”). Write a brief cover letter expressing your genuine interest in Amazon and why you’re excited about this opportunity in Vancouver. Double-check your application for spelling errors and ensure your contact information is current.

Interview Preparation: If selected for an interview, research Amazon‘s Leadership Principles, Edge AI initiatives, and engineering culture beforehand. Prepare specific examples using the STAR method (Situation, Task, Action, Result) to demonstrate your distributed systems design and technical leadership skills. Common questions may include scenarios about resolving ambiguous technical requirements, leading cross-team programs, and making architecture trade-off decisions under pressure. Dress appropriately for a technology environment, arrive 10–15 minutes early, and bring copies of your resume. Prepare thoughtful questions about the role, team structure, and platform roadmap. After the interview, send a thank-you email within 24 hours reiterating your interest in the position.