locaihost data

PDF to Markdown for RAG, with chunks ready to embed.

Send links to PDFs, Word files, slide decks, spreadsheets or web pages. Get clean Markdown back with headings and tables intact, or heading-aware chunks with page numbers, ready to embed and upsert into your vector database.

arxiv.org · 1706.03762 PDF

Attention Is All You Need

Pages
15
Words
5,697
Tokens, approx.
9,876
Chunks
34 at 512 tokens
Needs OCR
No
Billed pages
15
View the JSON record
{
  "sourceUrl": "https://arxiv.org/pdf/1706.03762",
  "fileName": "1706.03762v7.pdf",
  "format": "pdf",
  "title": "Attention Is All You Need",
  "pageCount": 15,
  "wordCount": 5697,
  "tokensApprox": 9876,
  "needsOcr": false,
  "chunkCount": 34,
  "chunks": [
    {
      "id": "dd93e90470f88607-2",
      "index": 2,
      "text": "## Abstract\n\nThe dominant sequence transduction models are based on complex recurrent or convolutional neural networks …",
      "tokensApprox": 288,
      "headingPath": ["Attention Is All You Need", "Abstract"],
      "page": 1,
      "pageEnd": 1
    }
  ],
  "billedPages": 15,
  "truncated": false,
  "error": null
}
A real record from a run on 10 Oct 2026, with the full Markdown and 33 other chunks left out. Your documents are processed for your run only.

Who it's for

RAG and AI agent builders
Feed a vector database from mixed PDFs, DOCX and web pages with one consistent output: chunks with heading paths and page numbers for citations.
Knowledge-base and search teams
Migrate manuals, policies and reports into Markdown that keeps its structure, so headings and tables still mean something after import.
Analysts and researchers
Pull tables out of PDFs and slide decks as Markdown tables, and get word and token counts to budget LLM calls before you make them.

How it works

Your documents
Public URLs, signed S3, GCS, Azure or Dropbox links, Google Docs exports, a file upload, or documents already in Apify storage. Nothing else is fetched.
Structure kept
Headings from font sizes and section numbers, tables rebuilt as Markdown tables, running headers, footers and page numbers removed.
Pay per page
PDF pages, slides and sheets count as pages; other formats count one page per 3,000 characters of output. Scanned PDFs are flagged and cost nothing.
Chunks for RAG
Choose chunk size and overlap in tokens. Each chunk carries its heading path and page range; one dataset item per chunk if you want it.

Pricing

Billed by Apify per result. Higher Apify plans pay less per result automatically.

Apify planPer 1,000 pages
Starter$1.00
Scale$0.75
Business$0.55

Per page converted, billed by Apify. A scanned PDF with no text layer returns needsOcr: true and is not charged. Use Max pages per document to cap cost on very long files.

API and integrations

Call it from any language over HTTP, or wire it up without code.

  • Schedules and saved tasks in the Apify console
  • Slack, email, Google Sheets, Zapier, Make, n8n and webhooks
  • The Apify MCP server, for AI agents
  • Python and JavaScript clients

Full API reference on Apify

curl -X POST \
  "https://api.apify.com/v2/acts/locaihost~doc-to-markdown/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "startUrls": [
      { "url": "https://arxiv.org/pdf/1706.03762" },
      { "url": "https://example.com/handbook.docx" }
    ],
    "outputMode": "chunks",
    "chunkSize": 512,
    "chunkOverlap": 64,
    "chunkContextHeader": true
  }'

Questions

Which formats are supported?

PDF, DOCX, PPTX (with speaker notes), XLSX (one table per sheet), CSV, HTML pages (main content only), Markdown and plain text. Mixed formats can go in one run and come back in one schema.

How is a page counted?

A PDF page, a PowerPoint slide or an Excel sheet is one page. Word, CSV, HTML, Markdown and text count one page per 3,000 characters of Markdown output, with a minimum of one, so a typical Word report comes out close to its printed length.

Does it do OCR on scanned PDFs?

Not in this version. A PDF with little or no text layer comes back with needsOcr: true and is not charged, so you never get an empty document that looks like it worked.

Are tables preserved?

Yes. Word, PowerPoint, Excel, CSV and HTML tables become GitHub-flavoured Markdown tables. PDF tables are rebuilt from text positions where columns line up, including tagged PDFs.

How do I load the chunks into a vector database?

Set output mode to chunks to get one dataset item per chunk with id, text, token estimate, heading path and page range. Send them to Pinecone, Qdrant, Weaviate or your own store through the Apify API or an integration.

Turn a folder of documents into clean chunks.

Paste document links, pick a chunk size, and get Markdown your retrieval pipeline can cite by page.