Skip to content
ToolsOnDuty - free online tools
PDF Tools

PDF to JSON

Convert PDF text and layout into structured JSON with pages, lines, cells and coordinates.

Free browser toolRuns in your browserNo sign-up

Structured PDF extraction studio

Reconstruct PDF text and layout as JSON

Preserve pages, lines, cells and source coordinates for integrations that need more structure than plain text.

Coordinates

layout data

Document hint

classification

Local

processing

Drop a PDF here, or click to choose
Text and layout coordinates stay in your browser.

About the PDF to JSON

PDF to JSON reconstructs selectable PDF text into pages, visual lines, and coordinate-aware cells. The output also records source information, page dimensions, statistics, a schema version, and a convenience document-type hint.

Coordinates preserve evidence that plain text loses, which is useful when an integration must rebuild columns or apply template-specific rules. The tool deliberately does not guess invoice totals, tax fields, or accounting meaning.

Image-only scans contain no readable text layer and must go through OCR first. Extraction and JSON generation run locally, but long coordinate-heavy documents can create large in-memory results.

Key features

  • Page, line, and cell hierarchy
  • Source coordinates and page dimensions
  • Document-type convenience hint
  • Word, character, line, and cell statistics
  • Pretty or compact JSON preview
  • Copy and JSON download actions
  • Versioned output schema

How to use

  1. 1Upload a PDF containing selectable text.
  2. 2Convert it and review the detected type and extraction statistics.
  3. 3Inspect page, line, cell, and coordinate structures in the preview.
  4. 4Toggle Pretty print depending on whether readability or compact size matters.
  5. 5Copy or download the JSON and validate it against the source document.

Examples

Prepare an invoice integration
Input: Selectable invoice with a line-item table
Output: JSON pages with lines, cells, and coordinates

Map values with vendor-specific rules and retain the source coordinates for review.

Frequently asked questions

Does it extract invoice fields automatically?
No. It preserves coordinate-aware source structure without silently assigning accounting fields.
Does it run OCR?
No. Image-only scans need OCR Searchable PDF first.
Why are coordinates included?
PDF text is positioned visually; coordinates help reconstruct rows and columns that plain text discards.
Is document-type detection authoritative?
No. It is a convenience hint and not business, tax, or accounting validation.
Does pretty printing change the data?
No. It changes JSON whitespace only.

Open a related tool to prepare your files or refine the finished result.