SKIP TO CONTENT
SD

SHREYAS DEB / PORTFOLIO

0LOADING
2026·AI / ML

Indic Language Models

Two fully independent, from-scratch decoder-only Transformer language models for Hindi and Nepali, built across three phases: data pipeline + tokenizer, pretraining + evaluation, and reasoning finetuning + attention analysis. No shared data, vocabulary, or weights between the two languages, and no pretrained models, tokenizers, or HuggingFace transformers classes anywhere; positional embeddings, multi-head attention, and causal masking are all implemented from scratch. Includes a custom BPE tokenizer swept across 5 vocab sizes, a KenLM-based perplexity filter for corpus cleaning, and a synthetic reasoning dataset generator for finetuning.

YEAR
2026
DOMAIN
AI / ML
STACK
PyTorchTransformersNLPDeep Learning