Skip to content
← Back to news
Researchanalysis20 months ago

Kathmandu University researchers build larger Nepali corpus and transformer models

Researchers from Kathmandu University's Information and Language Processing Research Lab reported a 27.5 GB Nepali corpus and Nepali BERT, RoBERTa and GPT-2 models. The peer-reviewed CHiPSAL 2025 paper reports a Nep-gLUE score of 95.60 and improvements over earlier Nepali systems.

ACL Anthology · AI Meridian research desk ·

Research snapshot verified 29 July 2026. News is independently written from cited primary and scholarly sources; provider claims remain attributed.

Key points

  • The corpus contains 27.5 GB of Nepali text.
  • The team trained BERT, RoBERTa and GPT-2 architectures for Nepali.
  • The work moved from a 2024 preprint to a peer-reviewed 2025 workshop paper.

Why it matters

A larger monolingual corpus and reproducible baseline models directly address the data scarcity constraining Nepali NLP.

Possible impact

Researchers can build stronger Nepali search, classification and generation systems, though compute, licensing and downstream evaluation remain constraints.

Important numbers

Nepali corpus
27.5 GB
Nep-gLUE
95.60
Model families
3

Evidence record

  1. Event or publication date — 24 November 2024
  2. Geography — Nepal; organisation — Kathmandu University — ILPRL
  3. Verified 2026-07-29 · peer reviewed cross checked · confidence 97%