• Projects
  • Contact

Summarizr

Turns a YouTube video's transcript into a concise summary, via a fine-tuned LED model behind a thin Go API.

GitHub
gopythonfastapipytorchterraformdocker

A pipeline that takes a YouTube video, pulls its transcript, and returns a concise summary — a thin Go API in front of a fine-tuned summarization model.

Architecture

A request crosses three hops: a Go API (gorilla/mux) validates the request and calls a Python FastAPI model service, which fetches the transcript and runs generation.

  • Transcript fetch prefers a manually-created caption track over an auto-generated one — auto-captions misspell names in a way that bleeds straight into the summary.
  • The model is a fine-tuned LED (Longformer Encoder-Decoder). Since LED only attends locally by default, the first token gets a global attention mask so it summarizes the input instead of just extending it. Input is capped at 8192 tokens (the fine-tune’s training window); output length scales with how much of that window is used.
  • Split across independently versioned repos — a superproject pinning the model, API, and web UI via git submodules — rather than one monorepo.

Training data

Fine-tuned first on QMSum, a public meeting-transcript dataset, to validate the training pipeline end to end before spending time scraping. A second phase — retraining on real scraped YouTube transcripts across several channels and genres — is currently blocked on a YouTube IP rate-limit ban clearing.

Status

The model and API work end to end locally. A GPU deployment (AWS EC2 via Terraform) is designed but not yet applied — CPU inference currently takes up to ~15 minutes per video, which is the whole reason that deployment exists. The web UI hasn’t been started yet, and Redis is wired into the Docker Compose setup but not yet used by the API.

The superproject is public; the model, API, and web UI repos it pins stay private while the project matures.