Learning Objectives

  • Integrate all previous modules into a working LM
  • Train a small model from scratch on real data
  • Understand practical training challenges
  • Generate text from the trained model

8.1 Project Specification

ComponentSpecification
Model size~40M parameters
Architecture8 transformer blocks, 512 hidden dim, 8 attention heads
DataShakespeare's complete works (~5MB)
TokenisationCharacter-level (vocab ≈ 80)
Context length256–512 tokens
Training time1–4 hours on modern GPU; longer on CPU
Batch size32–64

8.2 Data Preparation

Haskell — character-level tokeniser
-- Tokenise
tokenizeCharacter :: String -> (String -> [Int], [Int] -> String)
tokenizeCharacter text =
  let uniqueChars = nub (sort text)
      charMap     = zip uniqueChars [0..]
      idMap       = zip [0..] uniqueChars
      charToId c  = fromJust (lookup c charMap)
      idToChar i  = fromJust (lookup i idMap)
  in (map charToId, map idToChar)

data ShakespeareDataset = ShakespeareDataset
  { tokens    :: [Int]
  , charToId  :: Char -> Int
  , idToChar  :: Int  -> Char
  , vocabSize :: Int
  }

loadDataset :: FilePath -> IO ShakespeareDataset
loadDataset path = do
  text <- readFile path
  let (enc, dec) = tokenizeCharacter text
      uniqueChars = nub (sort text)
  return ShakespeareDataset
    { tokens    = enc text
    , charToId  = head . enc . (:[])
    , idToChar  = head . dec . (:[])
    , vocabSize = length uniqueChars
    }

8.3 Model Configuration

Haskell — Shakespeare model config
data Config = Config
  { vocabSize     :: Int
  , contextLength :: Int
  , d_model       :: Int
  , numHeads      :: Int
  , numLayers     :: Int
  , ffnDim        :: Int
  , dropout       :: Double
  }

shakespeareConfig :: Config
shakespeareConfig = Config
  { vocabSize     = 80    -- Character vocabulary
  , contextLength = 256
  , d_model       = 512
  , numHeads      = 8
  , numLayers     = 8
  , ffnDim        = 2048  -- 4× d_model
  , dropout       = 0.1
  }

-- Rough parameter count:
-- Embedding:       80 × 512         =  40,960
-- Per block:       ~2.1M params
-- 8 blocks:        ~16.8M params
-- Output proj:     512 × 80         =  40,960
-- Total:           ~17M–40M params

8.4 Complete Training Script

Haskell — main training entry point
main :: IO ()
main = do
  let config     = shakespeareConfig
      numEpochs  = 10
      batchSize  = 32

  -- Data
  dataset <- loadDataset "shakespeare.txt"
  let examples = createTrainingExamples (contextLength config) (tokens dataset)
      batches  = createBatches batchSize (contextLength config) examples

  -- Model
  model <- initializeModel config

  -- Training
  let initState = TrainingState
        { model     = model
        , optimizer = adamOptimizer
        , step      = 0
        , loss      = 0
        }

  putStrLn "Starting training..."
  (finalModel, finalState) <- trainLoop initState batches numEpochs config

  -- Save
  saveCheckpoint "checkpoints" finalModel (step finalState)

  -- Generate
  let prompt    = "To be"
      promptIds = map (charToId dataset) prompt
  generated <- generate finalModel promptIds 200 (idToChar dataset)
  putStrLn $ "Generated:\n" ++ generated

8.5 Text Generation

Haskell — autoregressive generation
-- Generate tokens autoregressively
generateTokens :: LanguageModel -> [Int] -> Int -> [Int]
generateTokens _     tokens 0             = tokens
generateTokens model tokens remainingLen  =
  let contextTokens    = takeContextWindow tokens
      logits           = forward model contextTokens
      nextTokenLogits  = last logits       -- last timestep
      nextToken        = sampleFromLogits nextTokenLogits
  in generateTokens model (tokens ++ [nextToken]) (remainingLen - 1)

-- Temperature sampling (lower = more deterministic, higher = more creative)
sampleWithTemperature :: Vector Double -> Double -> IO Int
sampleWithTemperature logits temperature = do
  let scaled = map (/ temperature) logits
      probs  = softmax scaled
  sampleCategorical probs

What to Expect

StageLossOutput quality
Random init~4.38 (ln 80)Random characters
1 epoch~3.0Mostly noise
3–5 epochs~2.0Learns spaces, punctuation
10–20 epochs~1.5Some word patterns visible
100+ epochs~0.8–1.0Shakespeare-like text

Example output (well-trained model)

ROMEO: O, gentle my lord, I would not be thy gentle love;

Reading & Viewing

Reference implementation

nanoGPT Andrej Karpathy ~400-line minimal GPT in PyTorch. Use as a reference to validate your Haskell implementation. github.com/karpathy/nanoGPT

Blog

The Unreasonable Effectiveness of Recurrent Neural Networks Andrej Karpathy Character-level generation fundamentals — still highly relevant. karpathy.github.io

Data

Shakespeare corpus (~5MB) Download and place in data/shakespeare.txt Direct download link