AI Researcher — Inference Optimization
Featherless AI
Hirezar Summary for South African Applicants
This fully remote full time position at Featherless AI is open worldwide, so South Africans can apply. This role is suited for senior-level professionals. As a remote position, you can work from anywhere in South Africa — whether you're based in Johannesburg, Cape Town, Durban, or a smaller town.
Job Description
Role Overview
We are seeking anAI Researcher with deep experience in inference optimizationto design, evaluate, and deploy high-performance inference systems for large-scale machine learning models. You will work at the intersection ofmodel architecture, systems engineering, and hardware-aware optimization, improving latency, throughput, and cost efficiency across real-world production environments.
Key Responsibilities
* Research and develop techniques tooptimize inference performancefor large neural networks.
Research and develop techniques tooptimize inference performancefor large neural networks.
* Improvelatency, throughput, memory efficiency, and cost per inference.
Improvelatency, throughput, memory efficiency, and cost per inference.
* Design and evaluatemodel-level optimizations(quantization, pruning, KV-cache optimization, architecture-aware simplifications).
Design and evaluatemodel-level optimizations(quantization, pruning, KV-cache optimization, architecture-aware simplifications).
* Implementsystems-level optimizations(dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization).
Implementsystems-level optimizations(dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization).
* Benchmark inference workloads across hardware accelerators.
Benchmark inference workloads across hardware accelerators.
* Collaborate with engineering teams todeploy optimized inference pipelines.
Collaborate with engineering teams todeploy optimized inference pipelines.
* Translate research insights intoproduction-ready improvements.
Translate research insights intoproduction-ready improvements.
Required Qualifications
* Strong background inmachine learning, deep learning, or AI systems.
Strong background inmachine learning, deep learning, or AI systems.
* Hands-on experience optimizing inference forlarge-scale models.
Hands-on experience optimizing inference forlarge-scale models.
* Proficiency inPythonand modern ML frameworks (e.g., PyTorch).
Proficiency inPythonand modern ML frameworks (e.g., PyTorch).
* Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime).
Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime).
* Ability to design experiments and communicate results clearly.
Ability to design experiments and communicate results clearly.
Preferred / Nice-to-Have Qualifications
* Experience deployingproduction inference systems at scale.
Experience deployingproduction inference systems at scale.
* Familiarity withdistributed and multi-GPU inference.
Familiarity withdistributed and multi-GPU inference.
* Experience contributing toopen-source ML or inference frameworks.
Experience contributing toopen-source ML or inference frameworks.
* Authorship or co-authorship of peer-reviewed research papersin machine learning, systems, or related fields.
Authorship or co-authorship of peer-reviewed research papersin machine learning, systems, or related fields.
* Experience working close to hardware (CUDA, ROCm, profiling tools).
Experience working close to hardware (CUDA, ROCm, profiling tools).
What Success Looks Like
* Measurable gains inlatency, throughput, and cost efficiency.
Measurable gains inlatency, throughput, and cost efficiency.
* Optimized inference systems running reliably in production.
Optimized inference systems running reliably in production.
* Research ideas successfully translated into deployable systems.
Research ideas successfully translated into deployable systems.
* Clear benchmarks and documentation that inform product decisions.
Clear benchmarks and documentation that inform product decisions.
Relevant Research Areas (Bonus)
* Long-context inference optimization
Long-context inference optimization
* Speculative decoding
Speculative decoding
* KV-cache compression and paging
KV-cache compression and paging
* Efficient decoding strategies
Efficient decoding strategies
* Hardware-aware inference design
Hardware-aware inference design
Originally posted onHimalayas
Tips for South African Applicants
Timezone Advantage
South Africa (SAST, UTC+2) overlaps well with European business hours and has a few hours of overlap with US East Coast. Mention your timezone flexibility in your application.
Salary in Context
Even without a listed salary, international remote roles typically pay 2-3x more than equivalent local positions in South Africa due to the exchange rate advantage.
Application Tips
Tailor your CV to international standards — use a clean format, highlight remote work experience, and include your English proficiency. Many SA applicants succeed by emphasising their strong work ethic and cultural adaptability.
Load Shedding Preparedness
If you're applying for a remote role, having a backup power solution (UPS, inverter, or generator) and mobile data as a backup internet connection shows employers you're prepared for South Africa's infrastructure challenges.
Related remote job searches
About Featherless AI
Featherless AI is a company in the Data Science industry that hires remote workers from South Africa. They currently have 4 open positions on Hirezar. View all Featherless AI jobs →