Skip to content
© Code by Kevin · Home
All projects

Case study · 02 / 033 min read

PDF Renamer

Automatic, content-based document classification

Role
Sole author
Team
Solo
Year
2026
Status
In daily use
Stack
JavaScript · HTML · CSS · PDF.js · OCR · File System Access API · Anthropic API

TL;DR

A document classifier that runs in the browser and tripled the school office's capacity: from ~100 to ~300 documents a day. Processing a class went from 3 days to 1 at most.

Context and problem

Students' digital records arrive as scanned PDFs with names like scan_0412.pdf. For each one the routine was the same: open it, read it, figure out the type, type the standard name, pick the student's folder and save. A class has about 200 documents, so organizing a whole class took 3 days.

Time wasn't the only problem. Hand-typed names drift ("Ficha medica", "FICHA MÉDICA", "ficha_saude"), and a file saved in the wrong folder disappears from the student's record.

My role

Sole author: I spotted the bottleneck while working in the office myself, designed the flow, built the tool and followed its adoption. The team uses it every day.

Constraints

  • No software installs: it had to run in the browser on the school's PCs.
  • Local files: records live in Windows folders, not on a server.
  • Scanned documents: many PDFs have no searchable text, just the image.
  • Non-developer users: document types and rules vary by department, so adjusting them couldn't depend on me.

Technical decisions and trade-offs

DecisionAlternativeWhy
Weighted keyword rules in an editable dictionaryA machine-learning modelThere was no labeled data to train on. Rules are explainable, and the office adjusts them from the UI.
Read the PDF's text first, OCR only when there's noneOCR every documentDigital PDFs are classified instantly; slower OCR is reserved for scans.
Read only the first 2 pagesThe whole documentThe type is almost always in the header, and in a batch every second counts.
A minimum confidence thresholdAlways suggest some typeWhen the score is low, the tool asks instead of being confidently wrong.
File System Access API writes straight into the student's folderDownload and move files by handRemoves two steps per document. Firefox falls back to a regular download.
Remember the folder with IndexedDBPick the folder every timeThe folder permission survives between sessions.

Architecture

One document's journey
  1. 01Batch of PDFsdrag and drop; queue with progress
  2. 02PDF.jsrenders and extracts the text
  3. 03Classifierrules + editable dictionary
  4. 04Suggested nametype + year, standardized
  5. 05Student's folderdirect write, duplicate check

Support

  • OCR (Tesseract.js)only for PDFs without a text layer
  • IndexedDB and localStoragefolder, profiles and settings
  • CSV historyeverything that was renamed

Implementation highlights

A classifier you can explain

The text is normalized (accents stripped, upper case) and each document type adds up the weights of the keywords it finds. Short terms must match a whole word, so "RG" (Brazil's ID card) isn't found inside another word. The winning type is only suggested above the threshold.

js/script.js (trimmed)
function scoreKeyword(text, kw) {
  // termos curtos exigem palavra inteira (evita "RG" casar dentro de outra palavra)
  if (kw.length <= 3) return new RegExp("\\b" + kw + "\\b").test(text);
  return text.includes(kw);
}
 
function classify(text) {
  const t = normalizeText(text);
  let best = null,
    bestScore = 0;
  for (const rule of CLASSIFY_RULES) {
    let score = 0;
    for (const [kw, w] of rule.kw) if (scoreKeyword(t, kw)) score += w;
    if (score > bestScore) {
      bestScore = score;
      best = rule.type;
    }
  }
  return bestScore >= CLASSIFY_THRESHOLD ? { type: best, score: bestScore } : null;
}

The built-in rules live alongside the ones the team creates in the UI: a department adds a keyword with a weight and it takes effect immediately, without me touching the code.

Built for daily use

  • Duplicates: if the name already exists in the folder, the tool asks what to do and suggests a free numbered name.
  • Keyboard: shortcuts to confirm, skip and move on, so hands never leave the keyboard mid-batch.
  • Exportable CSV history of everything that was renamed.

Results and impact

  • 3× more documents a day: from ~100 to ~300.
  • A class of ~200 documents went from 3 days to 1 at most.
  • Standard names and files in the right folder, which reduced filing errors across 545 students' records.

What I'd do differently and next steps

Links