SocialBlaze.ai

How to Do Cohort Analysis for Marketing (Step by Step)

How to Do Cohort Analysis for Marketing (Step by Step)

Table of Contents

Here’s the honest version of how to do cohort analysis: group your customers by when they arrived (or how they arrived), then watch each group separately over time instead of blending everyone into one big average. That’s it. That’s the whole technique. You build a simple table — one row per group, one column per month of their life with you — and suddenly you can see things your dashboard averages have been hiding for months: whether your onboarding change actually worked, which channels send people who stay, and whether your retention is quietly improving or quietly rotting underneath a flat-looking number.

Okay, let’s be honest about why this matters so much. Averages lie. Not maliciously — they just mix everything together. A flat overall retention number can hide improving cohorts drowned out by a flood of new signups, or decaying cohorts masked by aggressive acquisition. Cohorts are how you see change. And I promise: this is one of those techniques that sounds like it belongs to data scientists but is genuinely buildable in a spreadsheet by one marketer on a Tuesday afternoon.

Quick answer: how to do cohort analysis

  • Group people by a shared starting point — signup month, first-purchase month, campaign, or channel.
  • Track each group over its own timeline — month 0, month 1, month 2 of that group’s life, not calendar months.
  • Compare rows at the same age — March’s month-2 vs. January’s month-2 is the honest like-for-like.
  • Use time cohorts to see change and source cohorts to judge acquisition quality — the two moves that replace averages with truth.
  • Respect the limits — small cohorts are noisy, young cohorts are unfinished, and differences suggest causes rather than prove them.
Turn insight into a repeatable plan 1Audit your recentposts2Spot what alreadyworks3Make more of thewinners4Schedule itconsistently

What Is Cohort Analysis, Really?

Strip away the jargon and a cohort analysis is one move: group people by a shared starting point, then follow each group over time. A “cohort” is just everyone who shares that starting point — everyone who signed up in January, everyone who came from your spring campaign, everyone whose first purchase was a gift card.

Why bother? Because the alternative — the blended average — mixes your three-year veterans with the people who arrived yesterday and reports one number about the whole soup. If that number is flat, you learn almost nothing. Maybe nothing changed. Maybe your product got dramatically better for new users while your oldest users drifted away, and the two effects canceled out. Maybe retention is collapsing but a surge of new signups is propping the average up this month (new users are always “retained” at the start — they just got here).

Here’s the part nobody tells you: most “our metrics look fine” disasters are exactly this. The average looked fine. The cohorts underneath it did not. When you’re learning how to do cohort analysis, you’re really learning how to ask your data a sharper question: not “how are we doing?” but “how is each generation of customers doing, compared to the generations before them?”

How Do You Read a Cohort Table?

The classic cohort view is a retention grid, and once you can read one cell by cell, the whole technique clicks. Here’s a small illustrative example — every number here is fictional, invented purely to show the shape (please don’t treat anything in this grid as a benchmark):

Signup cohort Cohort size Month 0 Month 1 Month 2 Month 3 Month 4
January 400 100% 42% 31% 27% 25%
February 450 100% 44% 33% 29% —
March 520 100% 51% 40% — —
April 610 100% 54% — — —
May 700 100% — — — —

Illustrative retention grid with fictional numbers — the pattern is what matters, not the values.

Let’s walk it:

  • Each row is a cohort — everyone who signed up in that month. The row is that generation’s life story.
  • Each column is an age, not a date. “Month 1” means one month after that cohort’s signup — so January’s month 1 is February, but April’s month 1 is May. This is the single most important idea in the whole table.
  • Month 0 is always 100% — everyone is present at their own starting line. It’s the anchor, not an achievement.
  • Reading across a row shows how one generation decays over time. January kept 42% of its people at month 1, 31% at month 2, and the curve flattens near 25% — a settling pattern you’ll see often.
  • Reading down a column compares generations at the same age. Look at month 1: 42%, 44%, 51%, 54%. Something changed between February and March — maybe you shipped a better onboarding flow, maybe a better channel kicked in — and each new generation is sticking around more. Column comparison is your change detector.
  • The staircase of dashes is the diagonal — the edge of today. May hasn’t lived a month yet, so its row is short. The table is always a triangle: old cohorts have long rows, young cohorts have stubs.

Now here’s the kicker, and the reason cohort tables beat averages: in this fictional example, a blended “overall monthly retention” number could easily look flat or even declining across these months — because the fastest-growing cohorts (April, May) are huge and brand new, and new users always drag blended retention math around. Meanwhile the column view shows the truth: every new generation is doing better than the last. Same data. Opposite story. That’s the gap cohort analysis exists to close.

Time Cohorts vs. Source Cohorts: When They Joined vs. How

Most tutorials stop at time cohorts — grouping by signup or first-purchase month. Those answer “is the experience improving over time?” But for marketers specifically, the power move is the source cohort: grouping by how people arrived instead of when.

  • Time cohorts — by signup month, first-order month, trial-start week. Best for spotting change: did the thing we shipped in March make new customers stickier?
  • Source and behavior cohorts — by campaign, channel, landing page, plan, or first action taken. Best for judging quality: do customers from Channel A stay longer than customers from Channel B? Do people who complete a certain first step retain differently from people who don’t?

Source cohorts are where marketing budgets quietly get rescued. Two channels can deliver identical costs per signup while delivering wildly different customers — one sends people who stick around and buy again, the other sends people who vanish in week two. Averages can’t see that. A source cohort table makes it embarrassingly visible. If you only ever build one non-time cohort, make it this one.

What Questions Does Cohort Analysis Answer for Marketers?

A cohort table is only useful if you walk up to it holding a question. These are the five that earn the technique its keep.

Did our onboarding change actually work?

Compare the cohorts who arrived before the change to the cohorts who arrived after — at the same age. That last part is the honesty clause. The post-change cohort’s month-1 retention against the pre-change cohort’s month-1 retention is a fair fight. Comparing a two-month-old cohort’s current numbers to a one-year-old cohort’s current numbers is not — you’d be comparing a toddler to an adult and concluding the toddler is short. Same month-age, never same calendar month.

Which channels send customers who stay?

Build a source cohort: one row per acquisition channel, retention or repeat-purchase rate by month of customer age. This is the acquisition-quality truth serum. Cheap clicks that churn fast are expensive clicks wearing a disguise — you paid less per signup and more per customer who actually stayed. If churn is the leak you’re chasing, this pairs beautifully with a dedicated churn practice; my guide on how to measure churn covers the definitions and the denominator traps that make or break these comparisons.

Is our customer lifetime value actually improving?

Blended LTV drifts around with your customer mix and tells you very little. Cohort revenue curves — cumulative revenue per customer, by cohort, by month of age — are the only honest LTV trend. If each new cohort’s curve sits above the last cohort’s curve at the same age, value per customer is genuinely rising. If you want the full machinery, I’ve broken down how to measure customer lifetime value step by step, and cohort curves are the backbone of it.

Did the pricing or feature change help?

Same before/after logic as onboarding — with one piece of required honesty: other things changed too. You launched the new pricing in April, but April also brought a new campaign, a seasonal bump, and a competitor’s stumble. Before/after cohorts narrow the explanation; they don’t prove it. Cohort analysis tells you “the generations after the change behave differently.” It can’t tell you the change caused it. For proof you need a controlled test, which is its own discipline.

Are we comparing seasons fairly?

January signups and July signups are often different species — different intent, different context, different budgets. Comparing them raw and panicking about the “decline” is a classic self-inflicted wound. Cohort tables at least let you compare January-to-January, or ask whether this year’s January cohort beats last year’s at the same age. Stop comparing summer apples to New-Year’s-resolution oranges.

How to Do Cohort Analysis Step by Step (Even in a Spreadsheet)

Good news first: you may not need to build anything. Most serious analytics tools ship some form of retention or cohort view — GA4 has cohort exploration, and most product analytics and email platforms offer a version of it (features and names shift often, so verify what your current plan includes rather than trusting any article’s screenshot, including mine). If a built-in view answers your question, use it and spend your energy on interpretation.

But knowing how to do cohort analysis by hand, once, in a spreadsheet, is the best way to truly understand it — and for small datasets it’s genuinely practical. Here’s the walkthrough in plain words:

  • Step 1 — Export your raw list. You need, at minimum, one row per customer with two fields: when they started (signup date or first-purchase date) and evidence of activity over time (order dates, login months, or “active in month X” flags).
  • Step 2 — Assign each person a cohort. Add a column that converts their start date into a cohort label — usually the year-month, like 2026-03. For a source cohort, use the channel or campaign instead.
  • Step 3 — Compute each activity’s age. For every activity record, calculate the number of months between the cohort’s start month and the activity month. That’s the “month N” the activity belongs to. This step is the whole trick: you’re converting calendar time into cohort age.
  • Step 4 — Pivot. Build a pivot table: cohort labels as rows, age (month 0, 1, 2…) as columns, count of distinct active customers as the values.
  • Step 5 — Divide by cohort size. Convert the counts into percentages of each cohort’s month-0 size, so a 400-person cohort and a 700-person cohort become comparable.
  • Step 6 — Read columns, then rows. Columns tell you whether generations are improving. Rows tell you where in the lifecycle each generation leaks. Mark anything surprising and go find out why.

That’s the entire build. The first one takes an afternoon; every refresh after that takes minutes.

Which metric should the cells hold?

Retention is the default, but it’s not the law. Pick the cell metric that matches the business question:

  • Retention / activity rate — “do they keep showing up?” Best for subscriptions, apps, and communities.
  • Repeat-purchase rate — “do they buy again?” Best for ecommerce, where monthly logins mean little.
  • Revenue per cohort (cumulative) — “how much value does each generation produce?” The LTV-trend view.
  • Engagement actions — “do they do the thing that predicts staying?” Useful when revenue lags behavior by months.

One table, one metric. If you need two answers, build two tables — mixed-metric grids are how misreadings happen.

How Big Does a Cohort Need to Be?

Here’s the part that saves you from embarrassing conclusions: tiny cohorts are noise cosplaying as insight. If your March cohort is 12 people, then one person leaving moves the retention number by more than eight points. That’s not a trend; that’s Dave canceling his card. Don’t read drama into a 12-person month. There’s no universal minimum that applies to every business, but the working rule is simple: the smaller the cohort, the bigger the swing you should shrug off, and the more you should look for patterns that persist across several cohorts before believing them. If your monthly cohorts are tiny, group by quarter instead — fewer rows, steadier numbers, more truth.

The second honesty rule is maturity. Young cohorts have short rows — that triangle shape isn’t a flaw, it’s the edge of time itself. Resist the urge to judge a cohort’s month-6 destiny from its month-2 data. A two-month-old cohort hasn’t finished happening yet. Compare cohorts only at ages they’ve both actually reached, and label anything younger as “early read, subject to change.” The table will tempt you to extrapolate the missing cells. Don’t. Wait for the data to grow up.

How to Do Cohort Analysis Without Fooling Yourself

The table is easy. The reading is where people go wrong. Three disciplines keep you honest:

Watch for the improvement mirage (mix shift). A “better” recent cohort might not mean your product or onboarding improved — it might mean your channel mix changed. If April’s cohort is 60% referral traffic and January’s was 60% paid social, April isn’t a better-treated generation; it’s a different population. Before you celebrate a rising column, segment the cohorts by source and check whether the improvement survives. If it only exists in the blend, it’s a mix story, not an improvement story.

Keep correlation discipline. Cohort differences suggest causes; they don’t prove them. “The cohorts after the redesign retain better” is a lead, not a verdict — seasonality, pricing, mix, and luck all muddy the before/after water. When the stakes justify it, graduate from observation to experiment: a proper controlled test is the tool that turns “probably” into “demonstrated,” and it’s worth learning as a sibling skill to cohorts.

Mind the survivorship echo. The people still present in an old cohort’s month 12 are, by definition, the kind of people who stay — the leavers already left. So late-row behavior describes survivors, not the original population. “Month-12 customers love this feature” may really mean “the kind of people who stay a year love this feature.” Flip it around and the same echo warns you about research: if you only survey current customers, you’re only hearing from the rows’ right-hand edges.

What Should Your Cohort Ritual Look Like?

Cohort analysis works as a habit, not a heroic one-off. Two rhythms cover it:

  • A monthly glance (ten minutes). Refresh the table, read the newest column, compare it to the last few at the same age. You’re looking for one thing: is the newest generation on, above, or below the recent trend line?
  • A quarterly deep read (an hour). Rebuild the source-cohort view, check the LTV curves, and ask the five questions from earlier in this guide. This is where decisions get made — budget shifts, onboarding bets, channel pauses.

And one habit that future-you will thank present-you for: annotate the axis. Every time you launch something — a campaign, a pricing change, a redesign, a new channel — write it down against the cohort month it touched. Six months from now, when the March cohort’s column looks weird, you want a note that says “March: new onboarding flow + spring sale” instead of an archaeology expedition through old Slack threads. One line per change. That’s the whole habit, and it’s the difference between reading your table and excavating it.

Where does this fit in your wider measurement stack? Cohorts are a lens, not a KPI — they’re how you interrogate the KPIs you’ve already chosen. If your core metrics are still a bit of a junk drawer, start one level up with how to choose marketing KPIs, pick the handful that reflect real business health, and then point your cohort tables at exactly those.

Your Cohort Reading Checklist

Pin this next to the table. Before you act on anything a cohort grid tells you, run the list:

  • Same-age comparison? Am I comparing cohorts at the same month-age, not the same calendar month?
  • Cohort size? Are these cohorts big enough that one or two customers can’t swing the number?
  • Maturity? Am I judging a young cohort on cells it hasn’t lived yet?
  • Mix shift? Could a change in channel or audience mix explain this instead of a real improvement?
  • Annotation? What did we launch or change in this cohort’s early months?
  • Persistence? Does the pattern hold across several cohorts, or is it one odd row?
  • Causation humility? Am I treating this as a lead to investigate or a verdict to announce?

Questions-to-Cohorts Mapping Card

When you’re not sure which table to build, start from the question:

Your question Cohort type Cell metric What to compare
Did our onboarding change work? Time (signup month) Retention Pre-change vs. post-change cohorts at the same age
Which channels send keepers? Source (channel/campaign) Retention or repeat rate Channels against each other at the same age
Is LTV really improving? Time (first-purchase month) Cumulative revenue per customer Each cohort’s curve vs. older cohorts’ curves
Did pricing help or hurt? Time (before/after change) Revenue + retention Adjacent cohorts, with annotations for confounders
Is this seasonal or real? Time (same month, different years) Whatever’s in question January vs. last January at the same age
Does the first action predict staying? Behavior (did/didn’t do X) Retention Doers vs. non-doers from the same period

One more gentle plug, because it’s genuinely relevant: if social is one of your acquisition channels, the raw material for source cohorts is knowing which posts and platforms actually sent people. SocialBlaze’s analytics keep your per-platform performance in one place, so when you tag traffic by channel and build that source-cohort table, you can trace the winning rows back to the actual content that created them.

See which channels send customers who stay

SocialBlaze lets you schedule, auto-publish, and analyze every network from one dashboard — so the channel data feeding your cohort tables is clean, complete, and in one place. All on the Free Forever plan.

Start Free Forever →

FAQ: How to Do Cohort Analysis

What is cohort analysis in simple terms?

It’s grouping customers by a shared starting point — like their signup month or acquisition channel — and tracking each group separately over time. Instead of one blended average, you see how each generation of customers behaves, which reveals changes and differences that averages hide.

What’s the difference between a time cohort and a source cohort?

A time cohort groups people by when they arrived (signup month), which is best for spotting change over time. A source cohort groups them by how they arrived (channel, campaign, plan, or first action), which is best for judging acquisition quality — whether a channel sends customers who actually stay.

Can I do cohort analysis without special tools?

Yes. Export a customer list with start dates and activity dates, assign each person a cohort month, compute each activity’s age in months since that start, and pivot: cohorts as rows, ages as columns, active customers as values. Many analytics tools also include built-in cohort or retention views — check what your current plan offers.

How many people does a cohort need to be meaningful?

There’s no universal number, but the principle is firm: the smaller the cohort, the more one customer’s behavior swings the percentage, so the less any single cell means. If monthly cohorts are tiny, group by quarter, and only trust patterns that persist across several cohorts.

Does cohort analysis prove what caused a change?

No — it narrows the possibilities rather than proving a cause. Cohorts that behave differently after a change suggest the change mattered, but channel mix, seasonality, and other launches shift at the same time. Treat cohort differences as strong leads, and use controlled experiments when you need actual proof.

Frequently Asked Questions

Social Blaze provides a comprehensive suite of features including social media scheduling, analytics, content libraries, team collaboration tools, RSS feed automation, and a browser extension to streamline your social media strategy.

Absolutely! Social Blaze is designed to cater to both small businesses and larger agencies, offering customizable solutions to fit various needs, whether you’re managing a single account or multiple clients.

Our AI assistant takes the hassle out of content creation by creating AI post content for you, think of it as your social media sidekick, saving you time while helping you level up your strategy with smart insights.

Yes! Social Blaze offers various integrations with popular platforms and tools, allowing you to streamline your workflow and enhance your social media management experience seamlessly.

Table of Contents

×