ASR From Scratch
- Part 0 Prologue: The Questions 2026-07-19
- Part 1 What I'm trying to learn from my speech recognition experiments 2026-07-08
- Part 2 Speech recognition as independent parts: notes from a frozen-codebook ASR experiment 2026-07-08
- Part 3 Watching attention read speech: cluster-attention maps from a units-to-text transformer 2026-07-19
- Part 4 Why Attention Isn't the Problem: Understanding Classification Errors in Attention-Based Encoder-Decoder (AED) Models 2026-07-26
- Part 5 Beyond WER and CER: Why Indian-Language ASR Needs Akshara-Level, Phonetically-Weighted Metrics 2026-07-16
- Part 6 What speech is made of, and how much of it you need 2026-08-17
- Part 7 Filling the gaps: borrow the data, or record it? 2026-08-18
- Part 8 I downloaded data to fix the gaps. A control proved me half-wrong. 2026-08-18
TTS From Scratch
- Part 0 Prologue: The second set of questions 2026-08-09
- Part 1 Reading is not writing: the asymmetry at the heart of TTS 2026-08-10
- Part 2 What should a voice be made of? Two jobs, two kinds of token 2026-08-11
- Part 3 The wall: why you can't just predict the sound 2026-08-12
- Part 4 Say it, then voice it: splitting one hard job into two 2026-08-13
- Part 5 Coherent residuals: from independent heads to a depth transformer 2026-08-14
- Part 6 Adapt the consumer to the producer, again 2026-08-15
- Part 7 Where it fits: a map, and the voice it was aimed at 2026-08-16
The Shape of a Story
- Part 1 An embedding that remembers grammar 2026-08-02
- Part 2 Does a detective story have a shape? 2026-08-03
- Part 3 A language model with no words 2026-08-04
- Part 4 nanoGPT without a vocabulary 2026-08-05
- Part 5 Meaning is a mirage, grammar is real 2026-08-06
- Part 6 What it's actually good for 2026-08-07
- Part 7 A knob, not an engine 2026-08-09
- Part 8 A model that plans its own shape 2026-08-12
- Part 9 How much of it was just data? 2026-08-12
The Tiny Storyteller
- Part 1 A storyteller with ten thousand words 2026-08-25
- Part 2 Renting fluency: the same trick in Hindi and Telugu 2026-08-26
- Part 3 A knob for story shape 2026-08-26
- Part 4 Control tokens for slides 2026-08-26
Translation From Scratch
- Part 1 How machines learned to translate 2026-08-11
- Part 2 Translating structure, not words 2026-08-12
- Part 3 Indexing words, and the alignment problem 2026-08-13
- Part 4 Teaching a model Hindi agreement 2026-08-14
- Part 5 chrF lied, LaBSE didn't 2026-08-15
- Part 6 What structure-first translation is good for 2026-08-16
Varna
- Part 1 How far can a 93M phone recognizer go? Auditing varna against its 600M teacher across 12 Indian languages 2026-08-22
- Part 2 How India Speaks 2026-08-22
- Part 3 When the speaker disagrees with the label: accent phenomena inside an Indic phone recognizer 2026-08-28
Other writing
- 2026-08-09 Three things that fooled me building a codebook TTS
- 2026-08-08 Hand off the vector, not the id
- 2026-08-07 One codebook, both directions: a frozen alphabet as a contract
- 2026-08-06 Watching attention write speech: the cluster diagonal, in reverse
- 2026-07-08 Hello World