locaihost data

Broken link checker and sitemap audit, with redirects and canonicals.

One run reads the XML sitemap and crawls the site the way a search engine does. You get every URL with its status, redirect chain, canonical, noindex and hreflang, every broken link with the pages that link to it, and the orphan pages nobody links to.

keepachangelog.com 301 → 200

/en/1.1.0 redirects to /en/1.1.0/

Final URL
/en/1.1.0/
Redirect chain
1 hop · 301
Linked from
6 crawled pages
Canonical
Matches the final URL
Indexable
Yes
Response
47 ms · 26,341 bytes
View the JSON record
{
  "type": "url",
  "site": "https://keepachangelog.com",
  "url": "https://keepachangelog.com/en/1.1.0",
  "source": "crawl",
  "status": 200,
  "statusClass": "2xx",
  "redirectChain": [
    { "url": "https://keepachangelog.com/en/1.1.0", "status": 301 }
  ],
  "redirectHops": 1,
  "finalUrl": "https://keepachangelog.com/en/1.1.0/",
  "canonical": "https://keepachangelog.com/en/1.1.0/",
  "canonicalMismatch": false,
  "indexable": true,
  "hreflang": [],
  "hreflangIssues": [],
  "contentType": "text/html; charset=utf-8",
  "bytes": 26341,
  "responseTimeMs": 47,
  "title": "Keep a Changelog",
  "inlinks": 6,
  "outlinks": 14,
  "internalOutlinks": 1,
  "externalOutlinks": 13,
  "inSitemap": false,
  "charged": true
}
A real URL row from a run on 10 Oct 2026. Six pages link to the address without the trailing slash, so every one of those clicks pays a redirect hop. The site summary from the same run also flagged a homepage canonical pointing elsewhere and two pages with no canonical.

Who it's for

SEO agencies
Run it across every client site on a monthly schedule and send each client the broken-link and redirect list, without a desktop crawler licence.
Migrations and redesigns
Compare the old sitemap with what the new site actually serves. Find chains, loops and 404s before Google does.
Developers and international sites
Catch an accidental noindex or canonical regression after a deploy, and check that hreflang return links really match up across language versions.

How it works

Sitemap and crawl
Every Sitemap: line in robots.txt plus /sitemap.xml, sitemap indexes and gzipped files, compared against a breadth-first crawl from your start URL.
Polite by design
robots.txt is respected for its own user agent, one request at a time per site with a pause between. Bot walls (401, 403, 429) are marked unverified, not broken.
Pay per working URL
Charged only for URLs that end below status 400. Errors, broken-link rows, site summaries and external link checks are free.
Three views
One row per URL, one per broken link with up to 20 pages that link to it, and one summary per site, ready for CSV, Excel or JSON.

Pricing

Billed by Apify per URL that answers successfully. Higher Apify plans pay less per result automatically.

Apify planPer 1,000 URLs
Starter$1.00
Scale$0.75
Business$0.55

On the free Apify plan you can try it at the Starter price, up to 200 charged URLs per run. Set a total URL cap to limit the cost of any run.

API and integrations

Call it from any language over HTTP, or wire it up without code.

  • Schedules and saved tasks in the Apify console
  • Slack, email, Google Sheets, Zapier, Make, n8n and webhooks
  • The Apify MCP server, for AI agents
  • Python and JavaScript clients

Full API reference on Apify

curl -X POST \
  "https://api.apify.com/v2/acts/locaihost~url-inventory/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "sites": ["example.com", "https://shop.example.org/"],
    "mode": "both",
    "maxUrlsPerSite": 500,
    "checkExternalLinks": true
  }'

Questions

Sitemap, crawl or both?

Both is the default: it reads every sitemap listed in robots.txt plus /sitemap.xml, crawls from the start URL, and compares the two to find orphan pages. Sitemap mode checks only listed URLs; crawl mode suits sites without a sitemap.

Am I charged for broken links?

No. You pay only for URLs whose final status after redirects is below 400. 404s, 5xx errors, timeouts, DNS failures and redirect loops are reported for free, and so are external link checks and site summaries.

What about sites that block bots?

401, 403, 429 and LinkedIn's 999 are reported as unverified, not broken. If most of a site answers 403, the site summary carries a warning. Allow-list the user agent locaihost-url-inventory to audit it.

Are orphan pages reliable?

Orphan results are only reported when the crawl reached every linked page. If it stopped at your URL limit, the summary says so instead of guessing; raise Max URLs per site to finish the crawl.

Does it render JavaScript?

No. Pages are analysed as served, which is how most crawlers first see them.

Find every dead link before your visitors do.

Add your sites, pick a mode, add a monthly schedule. Broken links arrive with the pages you need to edit.