OUR SECTORS
At European Tech Recruit, our sectors cover a wide range of industries within the field of technology.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
At European Recruitment, our sectors cover a wide
range of industries within the field of technology
At European Recruitment, our sectors cover a wide
range of industries within the field of technology
Client services
Learn about the range of client services we offer at European Tech Recruit, and browse through our case sudies.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
About us
Learn about European Tech Recruit's mission, values, our team, and our commitment to DE&I.
At European Recruitment, our sectors cover a wide range of industries within the field of technology
Senior AI Operator Optimisation
Senior AI Operator Optimization Engineer — Contractor
Job Summary
We are seeking a highly motivated Senior AI Operator and Kernel Optimization Engineer to join an advanced AI infrastructure team.
In this role, you will optimize AI operators and kernels across a range of hardware architectures, including CPUs, GPUs, NPUs, AI accelerators, and heterogeneous computing platforms. Your work will support high-performance execution of large-scale AI training and inference workloads.
You will collaborate with software platform architects, compiler engineers, runtime developers, and AI researchers to improve the performance, scalability, and efficiency of next-generation AI computing systems.
Open-source contributions are considered a strong indicator of engineering expertise. Applicants are encouraged to provide links to GitHub or GitLab profiles, merged pull requests, technical blogs, publications, and other evidence of participation in open-source AI communities.
Responsibilities
-
Design, optimize, and accelerate AI operators and compute kernels across CPUs, GPUs, NPUs, AI accelerators, and heterogeneous platforms.
-
Analyze hardware and software bottlenecks to improve throughput, latency, memory utilization, and power efficiency.
-
Collaborate with platform architecture teams on hardware-software co-design for next-generation AI computing systems.
-
Develop architecture-aware optimization tools and AI agents for compute-intensive and memory-intensive deep learning operators.
-
Optimize communication operators used in distributed AI training and inference across large-scale computing clusters.
-
Improve collective communication operations, including AllReduce, AllGather, ReduceScatter, and Broadcast.
-
Optimize communication algorithms, topology awareness, data movement, and the overlap of communication with computation.
-
Work with compiler and runtime teams to support automatic operator optimization and hardware-specific code generation.
-
Benchmark, profile, and analyze AI workloads to identify optimization opportunities across the hardware-software stack.
-
Evaluate emerging AI hardware architectures, compiler technologies, runtime systems, and distributed computing techniques.
Required Qualifications
-
Master’s degree or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline.
-
Strong experience in computer architecture, processor architecture, hardware design, or hardware performance optimization.
-
Hands-on experience optimizing AI operators, numerical kernels, or high-performance computing workloads on modern hardware.
-
Strong understanding of memory hierarchies, cache systems, SIMD and SIMT execution, parallel programming, and data movement.
-
Experience with AI frameworks or inference platforms such as PyTorch, JAX, ONNX Runtime, or vLLM.
-
Proficiency in Triton, C or C++, and Python.
-
Experience with GPU programming environments or AI accelerator software stacks.
-
Strong benchmarking, profiling, performance analysis, and debugging skills.
Preferred Qualifications
-
Experience optimizing communication operators for distributed AI training or inference.
-
Experience with collective communication and networking libraries such as NCCL, RCCL, MPI, Gloo, UCX, or oneCCL.
-
Experience with compiler technologies such as MLIR, LLVM, TVM, XLA, Triton, or Apache IREE.
-
Experience developing for NPUs, TPUs, IPUs, wafer-scale systems, dataflow accelerators, or other specialist AI hardware.
-
Familiarity with distributed systems and high-performance interconnect technologies, including RDMA, InfiniBand, NVLink, PCIe, CXL, and high-speed Ethernet.
-
Publications or open-source contributions related to AI systems, compilers, high-performance computing, or distributed AI.
Open-Source Community Experience
Candidates with a strong record of contributing to the AI open-source ecosystem are highly valued.
Relevant experience may include:
-
Merged pull requests, accepted commits, performance improvements, feature implementations, or bug fixes in established open-source projects.
-
Publicly available evidence of technical contributions, including contribution histories, code reviews, design proposals, RFCs, release notes, or technical documentation.
-
Contributions to projects in areas such as deep learning frameworks, compilers, operator libraries, distributed communication, model serving, and AI infrastructure.
-
Experience as a maintainer, core contributor, reviewer, module owner, technical steering committee member, or community leader.
-
Participation in technical discussions, community governance, RFC reviews, roadmap development, or ecosystem initiatives.
Experience contributing to projects such as Triton, PyTorch, MLIR, LLVM, Hugging Face Transformers, JAX, NCCL, RCCL, or comparable ecosystems would be advantageous.
Additional Skills
-
Strong understanding of AI systems software, compilers, runtime systems, and hardware-software co-design.
-
Experience with AI workload performance modelling and bottleneck analysis.
-
Excellent analytical, debugging, and problem-solving skills.
-
Strong written and verbal communication skills.
-
Ability to collaborate effectively across architecture, compiler, runtime, systems, and research teams.
-
A strong interest in advancing the performance and efficiency of next-generation AI infrastructure.
Apply Now
By applying to this role, you acknowledge that we may collect, store, and process your personal data on our systems.
For more information, please refer to our
Privacy
Notice