Broken link checker and sitemap audit, with redirects and canonicals.
One run reads the XML sitemap and crawls the site the way a search engine does. You get every URL with its status, redirect chain, canonical, noindex and hreflang, every broken link with the pages that link to it, and the orphan pages nobody links to.
/en/1.1.0 redirects to /en/1.1.0/
- Final URL
- /en/1.1.0/
- Redirect chain
- 1 hop · 301
- Linked from
- 6 crawled pages
- Canonical
- Matches the final URL
- Indexable
- Yes
- Response
- 47 ms · 26,341 bytes
View the JSON record
{
"type": "url",
"site": "https://keepachangelog.com",
"url": "https://keepachangelog.com/en/1.1.0",
"source": "crawl",
"status": 200,
"statusClass": "2xx",
"redirectChain": [
{ "url": "https://keepachangelog.com/en/1.1.0", "status": 301 }
],
"redirectHops": 1,
"finalUrl": "https://keepachangelog.com/en/1.1.0/",
"canonical": "https://keepachangelog.com/en/1.1.0/",
"canonicalMismatch": false,
"indexable": true,
"hreflang": [],
"hreflangIssues": [],
"contentType": "text/html; charset=utf-8",
"bytes": 26341,
"responseTimeMs": 47,
"title": "Keep a Changelog",
"inlinks": 6,
"outlinks": 14,
"internalOutlinks": 1,
"externalOutlinks": 13,
"inSitemap": false,
"charged": true
}
Who it's for
- SEO agencies
- Run it across every client site on a monthly schedule and send each client the broken-link and redirect list, without a desktop crawler licence.
- Migrations and redesigns
- Compare the old sitemap with what the new site actually serves. Find chains, loops and 404s before Google does.
- Developers and international sites
- Catch an accidental noindex or canonical regression after a deploy, and check that hreflang return links really match up across language versions.
How it works
- Sitemap and crawl
- Every Sitemap: line in robots.txt plus /sitemap.xml, sitemap indexes and gzipped files, compared against a breadth-first crawl from your start URL.
- Polite by design
- robots.txt is respected for its own user agent, one request at a time per site with a pause between. Bot walls (401, 403, 429) are marked unverified, not broken.
- Pay per working URL
- Charged only for URLs that end below status 400. Errors, broken-link rows, site summaries and external link checks are free.
- Three views
- One row per URL, one per broken link with up to 20 pages that link to it, and one summary per site, ready for CSV, Excel or JSON.
Pricing
Billed by Apify per URL that answers successfully. Higher Apify plans pay less per result automatically.
| Apify plan | Per 1,000 URLs |
|---|---|
| Starter | $1.00 |
| Scale | $0.75 |
| Business | $0.55 |
On the free Apify plan you can try it at the Starter price, up to 200 charged URLs per run. Set a total URL cap to limit the cost of any run.
API and integrations
Call it from any language over HTTP, or wire it up without code.
- Schedules and saved tasks in the Apify console
- Slack, email, Google Sheets, Zapier, Make, n8n and webhooks
- The Apify MCP server, for AI agents
- Python and JavaScript clients
curl -X POST \
"https://api.apify.com/v2/acts/locaihost~url-inventory/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"sites": ["example.com", "https://shop.example.org/"],
"mode": "both",
"maxUrlsPerSite": 500,
"checkExternalLinks": true
}'
Questions
Sitemap, crawl or both?
Both is the default: it reads every sitemap listed in robots.txt plus /sitemap.xml, crawls from the start URL, and compares the two to find orphan pages. Sitemap mode checks only listed URLs; crawl mode suits sites without a sitemap.
Am I charged for broken links?
No. You pay only for URLs whose final status after redirects is below 400. 404s, 5xx errors, timeouts, DNS failures and redirect loops are reported for free, and so are external link checks and site summaries.
What about sites that block bots?
401, 403, 429 and LinkedIn's 999 are reported as unverified, not broken. If most of a site answers 403, the site summary carries a warning. Allow-list the user agent locaihost-url-inventory to audit it.
Are orphan pages reliable?
Orphan results are only reported when the crawl reached every linked page. If it stopped at your URL limit, the summary says so instead of guessing; raise Max URLs per site to finish the crawl.
Does it render JavaScript?
No. Pages are analysed as served, which is how most crawlers first see them.
Find every dead link before your visitors do.
Add your sites, pick a mode, add a monthly schedule. Broken links arrive with the pages you need to edit.