DESCRIPTION:
We're looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication libraries and frameworks like NCCL, NVSHMEM, and NIXL.
This is a ground-floor opportunity to work at the intersection of high-performance computing, networking, and machine learning infrastructure - building the systems that power the largest AI workloads in the cloud.
Key job responsibilities
- Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
- Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
- Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
- Design mechanisms to detect functional and performance regressions before they reach production
- Work across many instance types, software stacks, and Linux environments
BASIC QUALIFICATIONS:
- Bachelor's degree or above in Computer Science, Computer Engineering, or related fields
- Strong proficiency in C/C++
- Solid coursework or project experience in: 1/ Operating Systems (Linux internals, kernel concepts, memory management) 2/ Parallel Computer Architecture (multi-threading, SIMD, GPU programming, cache coherence) 3/ Distributed Systems (consensus, message passing, fault tolerance, scalability)
- Familiarity with Linux development environments and toolchains
PREFERRED QUALIFICATIONS:
- Internship experience in ML communications, HPC networking, or RDMA/high-speed interconnectsThe starting pay for this position is listed below. Starting Day 1 of employment, Amazon offers EAP, Mental Health Support, Medical Advice Line, 401(k) matching. Learn more about our benefits at https://hiring.amazon.com/why-amazon/benefits.