F

Machine Learning Engineer — Multilingual Data

Featherless AI

Fully Remote
Remote - Worldwide
Worldwide
SA / EMEA hours

Hirezar Summary for South African Applicants

This fully remote full time position at Featherless AI is open worldwide, so South Africans can apply. This role is suited for mid-level professionals. As a remote position, you can work from anywhere in South Africa — whether you're based in Johannesburg, Cape Town, Durban, or a smaller town.

Job Description

We’re looking for a Machine Learning Engineer to own and scale ourmultilingual data pipeline—from sourcing and curation to evaluation and continuous improvement. You’ll work closely with researchers and infra engineers to ensure our models perform robustly across languages, scripts, and cultural contexts.

This role sits at the intersection ofdata, research, and production MLand is ideal for someone who cares deeply about data quality, linguistic diversity, and model generalization beyond English.

What You’ll Do
* Design, build, and maintain large-scalemultilingual datasetsacross high- and low-resource languages
Design, build, and maintain large-scalemultilingual datasetsacross high- and low-resource languages
* Develop data pipelines forcollection, cleaning, normalization, deduplication, and labeling
Develop data pipelines forcollection, cleaning, normalization, deduplication, and labeling
* Implement quality filters using statistical, heuristic, and model-based methods
Implement quality filters using statistical, heuristic, and model-based methods
* Work with researchers to definelanguage coverage, benchmarks, and evaluation metrics
Work with researchers to definelanguage coverage, benchmarks, and evaluation metrics
* Analyze dataset bias, coverage gaps, and failure modes across regions and scripts
Analyze dataset bias, coverage gaps, and failure modes across regions and scripts
* Supporttraining, fine-tuning, and distillationworkflows with high-quality multilingual data
Supporttraining, fine-tuning, and distillationworkflows with high-quality multilingual data
* Continuously iterate on datasets based on model performance and real-world usage
Continuously iterate on datasets based on model performance and real-world usage

What We’re Looking For
* 3+ years of experience as an ML Engineer, Applied Scientist, or similar role
3+ years of experience as an ML Engineer, Applied Scientist, or similar role
* Strong experience working withmultilingual or non-English datasets
Strong experience working withmultilingual or non-English datasets
* Solid understanding of NLP fundamentals (tokenization, embeddings, language modeling)
Solid understanding of NLP fundamentals (tokenization, embeddings, language modeling)
* Experience building scalable data pipelines (Python, Spark, Ray, or similar)
Experience building scalable data pipelines (Python, Spark, Ray, or similar)
* Familiarity with Unicode, scripts, tokenization challenges, and language-specific quirks
Familiarity with Unicode, scripts, tokenization challenges, and language-specific quirks
* Comfort collaborating with researchers and translating research needs into production systems
Comfort collaborating with researchers and translating research needs into production systems

Nice to Have
* Experience withlow-resource languagesor multilingual benchmarks (e.g. FLORES, XTREME)
Experience withlow-resource languagesor multilingual benchmarks (e.g. FLORES, XTREME)
* Exposure to LLM training, fine-tuning, or distillation
Exposure to LLM training, fine-tuning, or distillation
* Linguistics background or experience working with native language experts
Linguistics background or experience working with native language experts
* Contributions to open-source datasets or ML tooling
Contributions to open-source datasets or ML tooling
* Experience with data quality evaluation at scale
Experience with data quality evaluation at scale

Why Join
* Real ownership over acore differentiatorof the product
Real ownership over acore differentiatorof the product
* Work on models used globally, not just in English-speaking markets
Work on models used globally, not just in English-speaking markets
* Small, high-caliber team with deep ML and systems experience
Small, high-caliber team with deep ML and systems experience
* Competitive compensation + meaningful equity at Series A stage
Competitive compensation + meaningful equity at Series A stage

Originally posted onHimalayas

Tips for South African Applicants

Timezone Advantage

South Africa (SAST, UTC+2) overlaps well with European business hours and has a few hours of overlap with US East Coast. Mention your timezone flexibility in your application.

Salary in Context

Even without a listed salary, international remote roles typically pay 2-3x more than equivalent local positions in South Africa due to the exchange rate advantage.

Application Tips

Tailor your CV to international standards — use a clean format, highlight remote work experience, and include your English proficiency. Many SA applicants succeed by emphasising their strong work ethic and cultural adaptability.

Load Shedding Preparedness

If you're applying for a remote role, having a backup power solution (UPS, inverter, or generator) and mobile data as a backup internet connection shows employers you're prepared for South Africa's infrastructure challenges.

About Featherless AI

Featherless AI is a company in the Data Science industry that hires remote workers from South Africa. They currently have 4 open positions on Hirezar. View all Featherless AI jobs →