OCR Turns a Page Into Text. A VLM Just Reads It.
OCR-then-parse converts a page to text outside the model, so a language model only ever sees what the OCR engine decided to keep. A vision-language model reads the page itself, pixels and text together, which is why it can answer a question about a specific box on the page while
