You need to agree to share your contact information to access this dataset

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this dataset content.

English-Kalenjin Dataset

This is an English-Kalenjin parallel text dataset prepared for machine translation research and model training.

The dataset combines mined English-Kalenjin pairs, manually translated synthetic English from Swahili-Kalenjin candidates, and a small direct manual collection set.

Access And Release Status

This dataset is intentionally kept gated for now.

  • Upstream ANV data appears gated.
  • Licensing and redistribution rights still need confirmation before public release.
  • A future open-source release should wait for licensing clarity and additional quality review.

If you are curious, researching Kalenjin, or interested in low-resource machine translation, feel free to request access. The gated status is mainly a temporary licensing and stewardship precaution, not a closed-door policy.

Dataset Summary

Total rows: 60664

Splits:

  • train.csv: 54598 rows
  • test.csv: 6066 rows

Columns

  • id: stable row identifier.
  • eng: English text.
  • kal: Kalenjin text.
  • domain: broad domain where available; otherwise unspecified.
  • dialect: Kalenjin dialect where available; otherwise unknown.
  • label: data construction label.
  • source_id: internal/provenance identifier where available.

Label Meanings

  • mined: English-Kalenjin pairs mined from ANV text data. These are useful but still review-needed.
  • synthetic: English side created by translating Swahili-Kalenjin candidates through a manual LLM-assisted batch workflow. The Kalenjin side was preserved from the candidate data.
  • manual_collection: direct manually collected English-Kalenjin entries.

Data Creation Process

  1. ANV CSV text files were mined for candidate English-Kalenjin and Swahili-Kalenjin pairs.
  2. Duplicate normalized pairs were removed globally.
  3. Swahili-Kalenjin candidates were exported in manual batches.
  4. English translations were produced in a model GUI and pasted back into batch output CSVs.
  5. Batch outputs were strictly verified for CSV shape, row counts, row IDs, and blank fields.
  6. Mined, synthetic, and manual collection rows were merged, deduplicated, shuffled, and split.

Quality Notes

This dataset is not fully human-reviewed. Known risks include:

  • Some mined rows may contain trimmed or imperfect Kalenjin text.
  • Some synthetic English translations may be imperfect.
  • The mined label should not be interpreted as gold-standard quality.
  • Kalenjin dialect labels are incomplete and may be unknown.

Intended Use

  • English-Kalenjin machine translation experiments.
  • Low-resource language data exploration.

Not Intended For

  • Treating every row as professionally reviewed ground truth.
  • Public redistribution until licensing and upstream permissions are clarified.
  • High-stakes use without additional review.

Citation

Citation information is not finalized. Licensing and attribution details should be completed before public release.

Downloads last month
61

Models trained or fine-tuned on mutaician/english-kalenjin-dataset

Space using mutaician/english-kalenjin-dataset 1