ALL PROJECTS01 / 18
2026·AI / ML
Indic Language Models
Two fully independent, from-scratch decoder-only Transformer language models for Hindi and Nepali, built across three phases: data pipeline + tokenizer, pretraining + evaluation, and reasoning finetuning + attention analysis. No shared data, vocabulary, or weights between the two languages, and no pretrained models, tokenizers, or HuggingFace transformers classes anywhere; positional embeddings, multi-head attention, and causal masking are all implemented from scratch. Includes a custom BPE tokenizer swept across 5 vocab sizes, a KenLM-based perplexity filter for corpus cleaning, and a synthetic reasoning dataset generator for finetuning.
- YEAR
- 2026
- DOMAIN
- AI / ML
- STACK
- PyTorchTransformersNLPDeep Learning