Machine Learning Engineer — Multilingual Data
Featherless AI
Hirezar Summary for South African Applicants
This fully remote full time position at Featherless AI is open worldwide, so South Africans can apply. This role is suited for mid-level professionals. As a remote position, you can work from anywhere in South Africa — whether you're based in Johannesburg, Cape Town, Durban, or a smaller town.
Job Description
We’re looking for a Machine Learning Engineer to own and scale ourmultilingual data pipeline—from sourcing and curation to evaluation and continuous improvement. You’ll work closely with researchers and infra engineers to ensure our models perform robustly across languages, scripts, and cultural contexts.
This role sits at the intersection ofdata, research, and production MLand is ideal for someone who cares deeply about data quality, linguistic diversity, and model generalization beyond English.
What You’ll Do
* Design, build, and maintain large-scalemultilingual datasetsacross high- and low-resource languages
Design, build, and maintain large-scalemultilingual datasetsacross high- and low-resource languages
* Develop data pipelines forcollection, cleaning, normalization, deduplication, and labeling
Develop data pipelines forcollection, cleaning, normalization, deduplication, and labeling
* Implement quality filters using statistical, heuristic, and model-based methods
Implement quality filters using statistical, heuristic, and model-based methods
* Work with researchers to definelanguage coverage, benchmarks, and evaluation metrics
Work with researchers to definelanguage coverage, benchmarks, and evaluation metrics
* Analyze dataset bias, coverage gaps, and failure modes across regions and scripts
Analyze dataset bias, coverage gaps, and failure modes across regions and scripts
* Supporttraining, fine-tuning, and distillationworkflows with high-quality multilingual data
Supporttraining, fine-tuning, and distillationworkflows with high-quality multilingual data
* Continuously iterate on datasets based on model performance and real-world usage
Continuously iterate on datasets based on model performance and real-world usage
What We’re Looking For
* 3+ years of experience as an ML Engineer, Applied Scientist, or similar role
3+ years of experience as an ML Engineer, Applied Scientist, or similar role
* Strong experience working withmultilingual or non-English datasets
Strong experience working withmultilingual or non-English datasets
* Solid understanding of NLP fundamentals (tokenization, embeddings, language modeling)
Solid understanding of NLP fundamentals (tokenization, embeddings, language modeling)
* Experience building scalable data pipelines (Python, Spark, Ray, or similar)
Experience building scalable data pipelines (Python, Spark, Ray, or similar)
* Familiarity with Unicode, scripts, tokenization challenges, and language-specific quirks
Familiarity with Unicode, scripts, tokenization challenges, and language-specific quirks
* Comfort collaborating with researchers and translating research needs into production systems
Comfort collaborating with researchers and translating research needs into production systems
Nice to Have
* Experience withlow-resource languagesor multilingual benchmarks (e.g. FLORES, XTREME)
Experience withlow-resource languagesor multilingual benchmarks (e.g. FLORES, XTREME)
* Exposure to LLM training, fine-tuning, or distillation
Exposure to LLM training, fine-tuning, or distillation
* Linguistics background or experience working with native language experts
Linguistics background or experience working with native language experts
* Contributions to open-source datasets or ML tooling
Contributions to open-source datasets or ML tooling
* Experience with data quality evaluation at scale
Experience with data quality evaluation at scale
Why Join
* Real ownership over acore differentiatorof the product
Real ownership over acore differentiatorof the product
* Work on models used globally, not just in English-speaking markets
Work on models used globally, not just in English-speaking markets
* Small, high-caliber team with deep ML and systems experience
Small, high-caliber team with deep ML and systems experience
* Competitive compensation + meaningful equity at Series A stage
Competitive compensation + meaningful equity at Series A stage
Originally posted onHimalayas
Tips for South African Applicants
Timezone Advantage
South Africa (SAST, UTC+2) overlaps well with European business hours and has a few hours of overlap with US East Coast. Mention your timezone flexibility in your application.
Salary in Context
Even without a listed salary, international remote roles typically pay 2-3x more than equivalent local positions in South Africa due to the exchange rate advantage.
Application Tips
Tailor your CV to international standards — use a clean format, highlight remote work experience, and include your English proficiency. Many SA applicants succeed by emphasising their strong work ethic and cultural adaptability.
Load Shedding Preparedness
If you're applying for a remote role, having a backup power solution (UPS, inverter, or generator) and mobile data as a backup internet connection shows employers you're prepared for South Africa's infrastructure challenges.
Related remote job searches
About Featherless AI
Featherless AI is a company in the Data Science industry that hires remote workers from South Africa. They currently have 4 open positions on Hirezar. View all Featherless AI jobs →