Datasets:
English-Kalenjin Dataset
This is an English-Kalenjin parallel text dataset prepared for machine translation research and model training.
The dataset combines mined English-Kalenjin pairs, manually translated synthetic English from Swahili-Kalenjin candidates, and a small direct manual collection set.
Access And Release Status
This dataset is intentionally kept gated for now.
- Upstream ANV data appears gated.
- Licensing and redistribution rights still need confirmation before public release.
- A future open-source release should wait for licensing clarity and additional quality review.
If you are curious, researching Kalenjin, or interested in low-resource machine translation, feel free to request access. The gated status is mainly a temporary licensing and stewardship precaution, not a closed-door policy.
Dataset Summary
Total rows: 60664
Splits:
train.csv: 54598 rowstest.csv: 6066 rows
Columns
id: stable row identifier.eng: English text.kal: Kalenjin text.domain: broad domain where available; otherwiseunspecified.dialect: Kalenjin dialect where available; otherwiseunknown.label: data construction label.source_id: internal/provenance identifier where available.
Label Meanings
mined: English-Kalenjin pairs mined from ANV text data. These are useful but still review-needed.synthetic: English side created by translating Swahili-Kalenjin candidates through a manual LLM-assisted batch workflow. The Kalenjin side was preserved from the candidate data.manual_collection: direct manually collected English-Kalenjin entries.
Data Creation Process
- ANV CSV text files were mined for candidate English-Kalenjin and Swahili-Kalenjin pairs.
- Duplicate normalized pairs were removed globally.
- Swahili-Kalenjin candidates were exported in manual batches.
- English translations were produced in a model GUI and pasted back into batch output CSVs.
- Batch outputs were strictly verified for CSV shape, row counts, row IDs, and blank fields.
- Mined, synthetic, and manual collection rows were merged, deduplicated, shuffled, and split.
Quality Notes
This dataset is not fully human-reviewed. Known risks include:
- Some mined rows may contain trimmed or imperfect Kalenjin text.
- Some synthetic English translations may be imperfect.
- The
minedlabel should not be interpreted as gold-standard quality. - Kalenjin dialect labels are incomplete and may be
unknown.
Intended Use
- English-Kalenjin machine translation experiments.
- Low-resource language data exploration.
Not Intended For
- Treating every row as professionally reviewed ground truth.
- Public redistribution until licensing and upstream permissions are clarified.
- High-stakes use without additional review.
Citation
Citation information is not finalized. Licensing and attribution details should be completed before public release.
- Downloads last month
- 61
Models trained or fine-tuned on mutaician/english-kalenjin-dataset
Translation • 0.6B • Updated • 38