PDF to Markdown for RAG, with chunks ready to embed.
Send links to PDFs, Word files, slide decks, spreadsheets or web pages. Get clean Markdown back with headings and tables intact, or heading-aware chunks with page numbers, ready to embed and upsert into your vector database.
Attention Is All You Need
- Pages
- 15
- Words
- 5,697
- Tokens, approx.
- 9,876
- Chunks
- 34 at 512 tokens
- Needs OCR
- No
- Billed pages
- 15
View the JSON record
{
"sourceUrl": "https://arxiv.org/pdf/1706.03762",
"fileName": "1706.03762v7.pdf",
"format": "pdf",
"title": "Attention Is All You Need",
"pageCount": 15,
"wordCount": 5697,
"tokensApprox": 9876,
"needsOcr": false,
"chunkCount": 34,
"chunks": [
{
"id": "dd93e90470f88607-2",
"index": 2,
"text": "## Abstract\n\nThe dominant sequence transduction models are based on complex recurrent or convolutional neural networks …",
"tokensApprox": 288,
"headingPath": ["Attention Is All You Need", "Abstract"],
"page": 1,
"pageEnd": 1
}
],
"billedPages": 15,
"truncated": false,
"error": null
}
Who it's for
- RAG and AI agent builders
- Feed a vector database from mixed PDFs, DOCX and web pages with one consistent output: chunks with heading paths and page numbers for citations.
- Knowledge-base and search teams
- Migrate manuals, policies and reports into Markdown that keeps its structure, so headings and tables still mean something after import.
- Analysts and researchers
- Pull tables out of PDFs and slide decks as Markdown tables, and get word and token counts to budget LLM calls before you make them.
How it works
- Your documents
- Public URLs, signed S3, GCS, Azure or Dropbox links, Google Docs exports, a file upload, or documents already in Apify storage. Nothing else is fetched.
- Structure kept
- Headings from font sizes and section numbers, tables rebuilt as Markdown tables, running headers, footers and page numbers removed.
- Pay per page
- PDF pages, slides and sheets count as pages; other formats count one page per 3,000 characters of output. Scanned PDFs are flagged and cost nothing.
- Chunks for RAG
- Choose chunk size and overlap in tokens. Each chunk carries its heading path and page range; one dataset item per chunk if you want it.
Pricing
Billed by Apify per result. Higher Apify plans pay less per result automatically.
| Apify plan | Per 1,000 pages |
|---|---|
| Starter | $1.00 |
| Scale | $0.75 |
| Business | $0.55 |
Per page converted, billed by Apify. A scanned PDF with no text layer returns needsOcr: true and is not charged. Use Max pages per document to cap cost on very long files.
API and integrations
Call it from any language over HTTP, or wire it up without code.
- Schedules and saved tasks in the Apify console
- Slack, email, Google Sheets, Zapier, Make, n8n and webhooks
- The Apify MCP server, for AI agents
- Python and JavaScript clients
curl -X POST \
"https://api.apify.com/v2/acts/locaihost~doc-to-markdown/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"startUrls": [
{ "url": "https://arxiv.org/pdf/1706.03762" },
{ "url": "https://example.com/handbook.docx" }
],
"outputMode": "chunks",
"chunkSize": 512,
"chunkOverlap": 64,
"chunkContextHeader": true
}'
Questions
Which formats are supported?
PDF, DOCX, PPTX (with speaker notes), XLSX (one table per sheet), CSV, HTML pages (main content only), Markdown and plain text. Mixed formats can go in one run and come back in one schema.
How is a page counted?
A PDF page, a PowerPoint slide or an Excel sheet is one page. Word, CSV, HTML, Markdown and text count one page per 3,000 characters of Markdown output, with a minimum of one, so a typical Word report comes out close to its printed length.
Does it do OCR on scanned PDFs?
Not in this version. A PDF with little or no text layer comes back with needsOcr: true and is not charged, so you never get an empty document that looks like it worked.
Are tables preserved?
Yes. Word, PowerPoint, Excel, CSV and HTML tables become GitHub-flavoured Markdown tables. PDF tables are rebuilt from text positions where columns line up, including tagged PDFs.
How do I load the chunks into a vector database?
Set output mode to chunks to get one dataset item per chunk with id, text, token estimate, heading path and page range. Send them to Pinecone, Qdrant, Weaviate or your own store through the Apify API or an integration.
Turn a folder of documents into clean chunks.
Paste document links, pick a chunk size, and get Markdown your retrieval pipeline can cite by page.