Audit BLEU vs. COMET on a Multilingual Customer-Support Corpus
Overview
What this challenge is about.
Compute BLEU, chrF++, and COMET on 600 multilingual support triples, then analyze correlations with human ratings. Earn a verifiable certificate.
The scenario
The startup (Series B, around 140 staff, around 22M monthly active users) currently makes ship/no-ship decisions partly on automatic-MT metrics and needs the right metric stack before scaling to 24 languages next year.
The Brief
What you'll do, and what you'll demonstrate.
Decide which automatic MT metric (or mix) the internal quality dashboard should standardize on, backed by per-language correlation with human judgement.
Earning criteria — what you'll demonstrate
- Apply lexical and learned MT metrics across multiple language pairs
- Quantify metric-to-human-judgement correlation
- Diagnose where individual metrics systematically fail
- Recommend a multi-metric dashboard stack with explicit reasoning
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Machine Translation
Master · Nlp
Strong alignment
This challenge maps to Machine Translation at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Mt Evaluation
Apply mt evaluation to solve real industry problems and demonstrate production-level capability.
- Neural Mt
Apply neural mt to solve real industry problems and demonstrate production-level capability.
- Statistical Analysis
Apply statistical analysis to solve real industry problems and demonstrate production-level capability.
- Multilingual Evaluation
Apply multilingual evaluation to solve real industry problems and demonstrate production-level capability.
- Benchmarking
Apply benchmarking to solve real industry problems and demonstrate production-level capability.
- Python
Write clean, efficient Python for data processing, automation, and backend services.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
Applied AI Scientist
Metric-stack audits across multiple languages and human judgements are the applied-AI-scientist's contribution to any multilingual ML org.
This challenge sharpens
- mt-evaluation
- statistical-analysis
- multilingual-evaluation
NLP Engineer
Knowing where each MT metric breaks per language is what makes NLP engineers credible on internationalization-heavy teams.
This challenge sharpens
- mt-evaluation
- neural-mt
- multilingual-evaluation
ML Researcher
Designing a fair metric-vs-human study with per-language correlations mirrors the rigor expected in MT-research projects.
This challenge sharpens
- statistical-analysis
- benchmarking
- mt-evaluation