SocialBlaze.ai

How to Create an XML Sitemap (and Keep It Honest)

How to Create an XML Sitemap (and Keep It Honest)

Table of Contents

Here’s how to create an XML sitemap, in one honest paragraph: let your platform do it. Nearly every modern CMS — WordPress, Shopify, Wix, Squarespace, and friends — either generates a sitemap automatically or does it with a well-known SEO plugin, usually at an address like yoursite.com/sitemap.xml. If you run a static site, a sitemap generator or a build-time plugin produces one in minutes. You only hand-write a sitemap if your site is tiny and never changes. Once it exists, you submit it in Google Search Console, reference it in your robots.txt file, and check the sitemaps report occasionally to catch problems early.

That’s the mechanical answer, and honestly, the mechanics are the easy part. The part that actually separates a helpful sitemap from a harmful one — and yes, a sloppy sitemap can quietly work against you — is knowing what a sitemap can and can’t do, what belongs in it, and how to keep it clean without babysitting it. So let’s walk through the whole thing together, from “what even is this file” to a hygiene checklist you can run in ten minutes.

Quick answer: how to create an XML sitemap

  • Use your CMS or an SEO plugin first. WordPress and most major platforms generate and update a sitemap automatically — that’s the right choice for almost everyone.
  • Only include pages you want indexed: canonical, indexable URLs that return a 200 status. Leave out redirects, noindexed pages, duplicates, and thin utility pages.
  • Keep lastmod honest. Only update it when the page meaningfully changes — Google has said it ignores lastmod from sites that fake it.
  • Submit it in Google Search Console and add a Sitemap: line to your robots.txt, then watch the discovered-vs-indexed gap in the sitemaps report.
  • A sitemap is a discovery aid, not a rankings lever. It helps crawlers find your pages; it doesn’t force indexing or boost positions.
Turn insight into a repeatable plan 1Audit your recentposts2Spot what alreadyworks3Make more of thewinners4Schedule itconsistently

What is an XML sitemap, really?

An XML sitemap is a machine-readable file — plain text, structured with XML tags — that lists the URLs on your site you’d like search engines to know about. For each URL it can also carry a little metadata, like when the page was last modified. Search engine crawlers fetch this file and use it as a map of your site: “here’s everything worth looking at, and here’s what’s changed lately.”

Now, let’s be honest about what that buys you, because this is where a lot of advice oversells. A sitemap is a discovery aid. It helps crawlers find pages they might otherwise miss or reach slowly — pages buried deep in your architecture, brand-new pages with no links pointing at them yet, pages orphaned by a redesign. That’s genuinely valuable, especially on larger or newer sites.

Here’s what a sitemap does not do, and I promise knowing this will save you frustration later:

  • It doesn’t force indexing. Listing a URL in your sitemap is a suggestion, not a command. Google still decides whether a page is worth indexing based on its quality, uniqueness, and signals from the rest of the web.
  • It doesn’t boost rankings. There’s no ranking credit for having a sitemap. A page found via sitemap ranks exactly as it would if it were found through a link.
  • It doesn’t override other signals. If a page is blocked in robots.txt or carries a noindex tag, putting it in the sitemap won’t rescue it — it just sends crawlers contradictory instructions.

So the honest framing is this: a sitemap makes sure search engines can find everything you care about, quickly and reliably. Whether those pages then earn indexing and rankings is a separate battle — one fought with content quality, internal linking, and overall site health. If you want the bigger picture of how all those crawl-and-index pieces fit together, my guide on how to do a technical SEO audit walks through the full system, and the sitemap is just one instrument in that orchestra.

Does your site actually need an XML sitemap?

Okay, let’s be honest — not every site gets the same value from a sitemap, and I’d rather tell you that than pretend it’s magic for everyone.

Sites that genuinely benefit:

  • Large sites. Hundreds or thousands of pages means crawlers can’t realistically rediscover everything on every visit. A sitemap tells them exactly what exists and what’s fresh.
  • New sites with few backlinks. Crawlers discover pages mostly by following links. A brand-new site with almost nothing pointing at it is hard to discover organically — the sitemap is your introduction.
  • Sites with deep or weak internal linking. If some pages sit five clicks from the homepage, or sections aren’t well cross-linked, pages can become effectively invisible. Same story for sites prone to orphan pages — product pages that fall out of category listings, old posts that nothing links to anymore.
  • Sites with rich media or fast-changing content. News publishers, video libraries, and image-heavy sites have specialized sitemap formats designed exactly for them (more on those below).

Sites that benefit less: a small site — say, under a few hundred pages — with clean internal linking, where every page is reachable within a couple of clicks, will usually get crawled just fine without a sitemap. Crawlers follow your navigation and find everything. If that’s you, a sitemap won’t hurt (and since your CMS probably makes one automatically, you may as well keep it), but it’s not the thing standing between you and better rankings. I’d rather you spend that energy on content.

Here’s the part nobody tells you, though: even on small sites, the sitemap earns its keep as a diagnostic tool. Once it’s submitted to Google Search Console, the sitemaps report shows you how many of your listed pages Google actually indexed — and that gap is one of the fastest health checks in SEO. So my honest advice: have one, keep it clean, and use it to monitor, even if discovery isn’t your bottleneck.

How do you create an XML sitemap?

This is the question you came for, so let’s take it tier by tier, from “you probably already have one” to “you’re writing it by hand, you beautiful minimalist.”

Option 1: your CMS generates it automatically (most people stop here)

If you’re on WordPress, you likely already have a sitemap. WordPress has included a basic native sitemap for years, and the popular SEO plugins generate a more configurable one — typically at /sitemap.xml or /sitemap_index.xml — that updates itself every time you publish, update, or delete content. The big hosted platforms (Shopify, Wix, Squarespace, and similar) generate sitemaps automatically too, usually with no setting to even toggle.

Your job in this scenario isn’t creation — it’s configuration. Open your sitemap in a browser and look at what’s in it. Most plugins let you exclude content types (tag archives, author pages, media attachment pages) that don’t belong in there. Ten minutes of pruning here is the highest-value sitemap work most site owners will ever do.

Option 2: a generator or build plugin for static sites

Running a static site generator — Hugo, Jekyll, Astro, Next.js, Eleventy, take your pick? Nearly every one has a built-in sitemap feature or a well-maintained plugin that produces the file at build time. This is the ideal setup, honestly: the sitemap regenerates on every deploy, so it can never drift out of date. Configure your exclusions (drafts, utility pages) once and forget it.

If your site isn’t built from a generator, online crawler-based sitemap generators exist: they crawl your site like a search engine would and output a sitemap file you upload. These work, but notice the trap — the file is a snapshot. The moment you add a page, it’s stale. If you go this route, you’re signing up to regenerate it on a schedule, and in my experience, “I’ll remember to re-run it” is a promise future-you rarely keeps. Prefer anything automated.

Option 3: writing one by hand (tiny sites only)

For a genuinely tiny site — a handful of pages that change a few times a year — hand-rolling a sitemap is fine. The format is refreshingly simple: an XML declaration, a urlset wrapper, and one url entry per page containing a loc tag with the full absolute URL. Add a lastmod date if you’ll actually maintain it honestly; skip it if you won’t. Save the file as sitemap.xml at your site’s root, and you’re done.

Two rules if you hand-roll: use full absolute URLs (https and all, matching your canonical version exactly — www or non-www, trailing slash or not), and remember that every edit to your site now comes with the chore of editing this file too. The moment that chore starts slipping, switch to an automated option.

What belongs in your XML sitemap — and what doesn’t?

Here’s the mental model I want you to tattoo somewhere handy: your sitemap is a curated list of the pages you want indexed, not an inventory of every URL your site can produce. Search engines treat the sitemap as a statement of what you consider important. When that statement is clean and accurate, crawlers learn to trust it. When it’s full of junk, you’re sending mixed signals — “please index this!” next to pages that say “don’t index me” — and that inconsistency makes every signal you send a little less credible.

What belongs in:

  • Canonical URLs only. If a page has multiple versions (parameters, duplicates, print views), list only the canonical one — the version your canonical tags point to.
  • Indexable pages. Pages without a noindex tag, not blocked by robots.txt.
  • Pages returning a 200 status. Live, working pages — not redirects, not errors.
  • Pages you actually want ranking. Your real content: service pages, products, articles, landing pages.

What stays out:

  • Noindexed pages. “Index this” (sitemap) plus “don’t index this” (meta tag) is a contradiction. The noindex wins, and the contradiction erodes trust in your sitemap.
  • Redirecting URLs. List the destination, not the redirect. A sitemap full of 301s makes crawlers do pointless extra hops.
  • Duplicate and parameterized URLs. Filtered views, session IDs, UTM-tagged links — none of it.
  • Thin and utility pages. Cart, login, thank-you pages, internal search results, most tag archives. These exist for users mid-task, not for search.
  • Broken pages. Anything returning a 404 or 500 needs fixing or removing, not a sitemap listing. If your Search Console is already flagging these, my walkthrough on how to fix crawl errors covers triaging them properly.

A dirty sitemap won’t tank your site overnight — let’s keep perspective. But it wastes crawl attention on dead ends, muddies your diagnostics (you can’t read the indexed-vs-submitted gap when half your submissions were never index-worthy), and teaches crawlers your sitemap is noise. Curate it like you mean it.

What do the sitemap tags actually mean?

Time for a plain-words tour of the format, because the tags are simple but two of them come with honesty tests.

loc — the page’s full, absolute URL. This is the only required tag per entry, and the only one that’s unambiguously important. Match your canonical URLs exactly: protocol, subdomain, trailing slash, everything.

lastmod — the date (optionally with a time) the page last meaningfully changed. This tag is genuinely useful: it helps crawlers prioritize revisiting pages that have actually changed instead of re-fetching everything. Google has said it uses lastmod when it’s consistently accurate. And that’s the catch — keep it honest. Some plugins and generators stamp every URL with today’s date on every rebuild, which claims your entire site changed today. Do that, and Google learns your lastmod is meaningless and starts ignoring it — you’ve burned the one piece of sitemap metadata that actually works. Fake freshness isn’t a growth hack; it’s a credibility tax. Only update lastmod when the content meaningfully changes: real edits, new sections, corrected information. Not a footer year change, not a rebuild timestamp.

priority and changefreq — you’ll see these in older tutorials, where priority supposedly tells Google how important a page is (0.0 to 1.0) and changefreq how often it changes. Honest status check: Google has said for years that it largely ignores both, because site owners — shocker — set every page to priority 1.0 and changefreq “daily.” Most modern generators omit them or fill them with defaults nobody reads. Don’t spend a minute tuning them, and if you’re reading this well after I wrote it, verify the current guidance in Google’s own sitemap documentation — this stuff does evolve, and the docs are the source of truth.

One more format note: sitemaps must be UTF-8 encoded, and special characters in URLs (like ampersands) need escaping. Any decent generator handles this; it mostly bites the hand-rollers.

When do you need a sitemap index file?

A single sitemap file has hard limits: 50,000 URLs or 50MB uncompressed, whichever comes first. Cross either line and search engines may not process the file fully. (Verify those numbers against current documentation if you’re reading this down the road — they’ve been stable for years, but “stable so far” isn’t “guaranteed forever.”)

The solution is a sitemap index: a parent file that lists child sitemap files instead of pages. Think of it as a table of contents — sitemap-posts-1.xml, sitemap-products.xml, sitemap-pages.xml — each child staying under the limits. You submit the index file once, and search engines follow it to all the children.

Here’s the nice surprise: if you’re on WordPress with an SEO plugin, you already have one. That /sitemap_index.xml URL splitting your content by type? That’s a sitemap index. Beyond staying under limits, splitting by content type is quietly brilliant for diagnostics — when Search Console shows indexing stats per child sitemap, you can see at a glance that your posts index beautifully while your product pages struggle. That’s a clue you’d never get from one giant undifferentiated file. Larger sites sometimes split even further (by category, or by date) for exactly this reason.

What about image, video, and news sitemaps?

Quick tour of the specialized variants, because most sites need at most one of them:

  • Image sitemaps add image information to your URL entries — which images live on which pages. Useful for sites where image search traffic genuinely matters (photography, products, recipes, portfolios). For everyone else, regular crawling of your pages finds your images fine.
  • Video sitemaps carry video metadata — title, description, duration, thumbnail — and help your videos appear in video search results. Worth it if video is hosted on your pages and is a real part of your strategy.
  • News sitemaps are a specialized format for publishers in Google News, listing only articles from the last 48 hours. If you’re not a news publisher, this one isn’t for you.

My honest advice: don’t build these speculatively. Add a variant when the corresponding traffic channel is actually part of your plan. A standard sitemap plus good on-page markup covers the vast majority of sites, and several major SEO plugins can generate the variants when you do need them.

How do you submit your sitemap to Google?

Creating the file is half the job; now you tell search engines where it lives. Two moves, both quick:

1. Submit it in Google Search Console. Open your property, head to the Sitemaps report (under Indexing), paste in your sitemap’s path — sitemap.xml or sitemap_index.xml — and hit submit. If you use a sitemap index, submit just the index; Google follows it to every child. Bing has an equivalent flow in Bing Webmaster Tools, and it’s worth the extra two minutes. Submission isn’t a one-time ping, either — once submitted, Google re-fetches your sitemap periodically on its own, which is why automated freshness (coming up) matters more than re-submitting manually.

2. Reference it in your robots.txt. Add a single line to the robots.txt file at your site root: Sitemap: followed by the full sitemap URL. This makes your sitemap discoverable by any crawler that reads robots.txt — not just the ones you submitted to — and it acts as a safety net if a Search Console submission is ever lost or forgotten. If robots.txt is new territory, my guide on how to use robots.txt covers the whole file, including this line and the crawl rules that need to play nicely with your sitemap.

That’s it. No tricks, no paid fast lanes. And a gentle reality check while we’re here: submitting a sitemap gets your pages seen faster — it does not get them indexed or ranked faster than they deserve. Anyone promising otherwise is selling something.

How do you monitor your sitemap after submitting it?

Here’s where the sitemap turns from plumbing into a genuinely useful instrument. The Search Console sitemaps report shows, for each submitted sitemap, how many URLs were discovered and how many ended up indexed. Read that gap like a doctor reads a chart:

  • Small gap (most pages indexed): healthy. Some gap is normal — paginated archives and marginal pages often don’t make the cut, and that’s fine.
  • Large gap (many discovered, few indexed): diagnostic gold. Google found your pages and chose not to index them. Click through to the page indexing report and read the reasons: “Crawled — currently not indexed” usually points to quality or thin-content issues; “Duplicate without user-selected canonical” points to canonicalization problems; “Excluded by noindex” means noindexed pages are sitting in your sitemap and your curation needs work.
  • Errors on the sitemap itself: “Couldn’t fetch” or parse errors mean the file is unreachable or malformed — fix the file first, then worry about its contents.

Make this a monthly two-minute ritual: open the report, scan the gap, investigate anything that moved sharply. A sudden drop in indexed count is often your earliest warning that something broke — a deploy that accidentally noindexed a template, a plugin update that mangled the sitemap, a section of the site that fell over. The sitemap report catches these days before your traffic chart does.

How do you keep your sitemap automatically fresh?

A sitemap is only useful while it’s true. Every page you publish, update, or delete changes what the file should say — so the single most important architectural decision is this: your sitemap must regenerate automatically.

If your CMS or SEO plugin generates it dynamically, you have this for free — publish a post, and it appears in the sitemap seconds later; delete one, and it vanishes. Static-site builds regenerate on every deploy, which is just as good. The setup to avoid is the generate-once-and-upload pattern: a crawler-generated file from eight months ago that still lists deleted pages and knows nothing about new ones. A stale sitemap is worse than none, because it actively misdirects crawlers toward 404s while hiding your new content.

Quick annual audit question, and be honest with yourself: “If I published a page right now and touched nothing else, would it show up in my sitemap today?” If the answer is no, fix the pipeline before you fix anything else in this article.

What are the most common XML sitemap mistakes?

Let me save you from the greatest hits — these are the issues that show up over and over in audits:

  • Listing URLs that robots.txt blocks. You’re simultaneously saying “index this” and “you may not look at this.” The crawler can’t fetch the page, so it can’t index it properly — and the contradiction is a classic audit red flag.
  • Listing noindexed URLs. Same contradiction, different mechanism. Usually caused by a plugin including a content type you noindexed elsewhere.
  • Listing redirects and 404s. Sitemaps full of dead and moved URLs waste crawl attention and pollute your diagnostics.
  • The stale snapshot. The generated-once file that no longer reflects the site. Covered above, worth repeating — this is the most common sitemap failure on small-business sites.
  • Fake lastmod. Every URL stamped with today’s date, every day. Google learns to ignore your lastmod entirely, and you lose the tag’s real value.
  • Ignoring the size limits. A single file quietly crossing 50,000 URLs, with everything past the line at risk of being unprocessed. Big sites: use an index.
  • Wrong URL version. Sitemap lists http:// while the site is https://, or www while the canonical is non-www. Every entry triggers a redirect hop. Match your canonicals exactly.
  • Multiple conflicting sitemaps. An old plugin’s sitemap and a new plugin’s sitemap both live, listing different things. Pick one generator; retire the rest.

What should your sitemap hygiene checklist include?

Here’s the ten-minute checkup. Run it quarterly, or after any major site change — a redesign, a migration, a new plugin:

  • ☐ The sitemap loads. Open the URL in a browser. You should see XML, not a 404 or an error page.
  • ☐ It regenerates automatically. Publish-and-check test: a new page appears without manual work.
  • ☐ Spot-check 10 listed URLs. Each returns 200, is the canonical version, and has no noindex tag.
  • ☐ No junk content types. No media attachment pages, internal search results, cart/login/thank-you pages, or redundant archives.
  • ☐ lastmod is honest. Dates vary across URLs and reflect real edits — not one identical timestamp everywhere.
  • ☐ URL versions match canonicals. Protocol, subdomain, and trailing-slash style are consistent with your canonical tags.
  • ☐ Under the limits. No single file over 50,000 URLs or 50MB; an index file splits things up if needed.
  • ☐ Submitted and referenced. Present in Search Console (and Bing Webmaster Tools) with no fetch errors; a Sitemap: line exists in robots.txt.
  • ☐ The indexed gap is stable. Discovered-vs-indexed in the sitemaps report hasn’t moved sharply since last check.
  • ☐ Only one sitemap system is live. No orphaned sitemap from a retired plugin still floating around.

Troubleshooting: common sitemap problems and fixes

Symptom Likely cause Fix
Search Console says “Couldn’t fetch” Wrong URL submitted, sitemap blocked by robots.txt, or server error Open the sitemap URL in a browser; correct the path, unblock it in robots.txt, or fix the server issue, then resubmit
“Sitemap could not be read” / parse error Malformed XML — unescaped characters, wrong encoding, or an HTML error page served at the sitemap URL Validate the XML; let a plugin or generator produce the file instead of hand-editing
Many URLs discovered, few indexed Quality, duplication, or canonicalization issues — or junk URLs padding the sitemap Read the reasons in the page indexing report; prune non-index-worthy URLs; improve thin content; fix canonicals
“Excluded by noindex” on sitemap URLs Sitemap includes pages marked noindex Remove those URLs from the sitemap, or remove the noindex if the exclusion was accidental
Indexed count dropped sharply A deploy or plugin change broke templates (accidental noindex), or a site section is erroring Compare recent changes to the drop date; inspect affected URLs with the URL inspection tool
New pages take ages to appear in the sitemap Static or cached sitemap not regenerating Switch to dynamic generation or regenerate on every deploy; clear aggressive caching on the sitemap URL
Google reports more sitemaps than you submitted An old plugin’s sitemap still live alongside the current one Disable the retired generator; remove stale sitemaps from Search Console; keep one source of truth

Where does this fit in the bigger picture?

I’ll leave you with the framing that keeps sitemap work proportional. Your sitemap is the invitation list; your content is the party. Learning how to create an XML sitemap — and keeping it clean, honest, and automatic — ensures every page you care about gets found promptly and gives you an early-warning dashboard for indexing health. That’s real value. But it’s enabling infrastructure, not a growth strategy: set it up properly once, check it quarterly, and pour the hours you saved into the pages themselves and the links between them. The sites that win aren’t the ones with the fanciest sitemaps — they’re the ones whose sitemaps list pages worth indexing.

Spend your time on content, not plumbing

Once your sitemap quietly does its job, the growth comes from publishing and promoting consistently. SocialBlaze lets you schedule, auto-publish, and analyze your content across every social network from one place — so each new page gets an audience, not just an index entry. Start on the Free Forever plan.

Start Free Forever →

FAQ: how to create an XML sitemap

Do I need an XML sitemap if my site is small?

Probably not for discovery — a small site with clean internal linking gets crawled fine without one. But since most platforms generate a sitemap automatically anyway, keep it: submitted to Search Console, it doubles as a free indexing-health monitor, and the discovered-vs-indexed gap is a diagnostic you’ll be glad to have.

Does an XML sitemap improve rankings?

No. A sitemap is a discovery aid: it helps search engines find your pages faster and more reliably, especially on large or new sites. It doesn’t force indexing and carries no ranking boost. Indexing and rankings are earned by page quality, relevance, and links — the sitemap just makes sure nothing worth ranking goes unseen.

How often should I update my XML sitemap?

Continuously and automatically. Your sitemap should regenerate whenever you publish, update, or delete content — which CMS plugins and build-time generators do for you. If yours requires manual regeneration, that’s the thing to fix; a stale sitemap misdirects crawlers toward deleted pages while hiding new ones.

Should every page on my site be in the sitemap?

No — only canonical, indexable pages that return a 200 status and that you actually want in search results. Leave out redirects, noindexed pages, duplicates, and utility pages like carts, logins, and thank-you pages. A curated sitemap sends consistent signals; a dump-everything sitemap sends contradictions.

What’s the difference between an XML sitemap and an HTML sitemap?

An XML sitemap is a machine-readable file built for search engine crawlers — people never see it. An HTML sitemap is a regular web page listing your site’s sections for human visitors, which also adds some internal links. They serve different audiences; the XML version is the one search engines expect you to submit.

Frequently Asked Questions

Social Blaze provides a comprehensive suite of features including social media scheduling, analytics, content libraries, team collaboration tools, RSS feed automation, and a browser extension to streamline your social media strategy.

Absolutely! Social Blaze is designed to cater to both small businesses and larger agencies, offering customizable solutions to fit various needs, whether you’re managing a single account or multiple clients.

Our AI assistant takes the hassle out of content creation by creating AI post content for you, think of it as your social media sidekick, saving you time while helping you level up your strategy with smart insights.

Yes! Social Blaze offers various integrations with popular platforms and tools, allowing you to streamline your workflow and enhance your social media management experience seamlessly.

Table of Contents

×