How well do modern language models work for African languages with little written data, and what breaks first?
- Approach
- Build evaluation sets in named languages, publish the failure modes, and test whether smaller task-specific models beat large general ones on local tasks.
- What counts as an output
- Open evaluation sets, reproducible benchmarks, and a documented cost floor for running these systems on regional infrastructure.

