Retrospectives
A team that never stops to examine how it works will keep repeating the same mistakes. Retrospectives turn that examination into concrete, committed change.
Overview
A retrospective is not a ceremony. It is the reflect beat in a loop your team is already running, whether or not anyone has named it:
Plan. Commit to an outcome and a method for getting there. Do. Build the thing. Share. Put the work in front of stakeholders and watch what actually happens — the demo, the release, the customer call. Reflect. Process what came back, and change the plan.
Skip the last two beats and you get a team that ships on schedule and learns nothing. The retrospective is where feedback stops being an inbox and becomes a decision. Run well, it feeds directly into better flow by surfacing the friction that slows delivery, and it reinforces small slices of value thinking — teams experiment with incremental changes to how they work instead of announcing a new process every quarter.
The tell for a working retro is not how the hour felt. It’s whether the plan is different when you leave.
Three Retrospectives Worth Running
The default retro — three columns, sticky notes, twelve items nobody will look at again — is what teams reach for when nobody chose a format. Choose one. Each of these produces a decision, not a list.
The One Where We Interpret Metrics
Run it when the team’s numbers have moved and everyone has a theory. Throughput sagging, a widening band on the cumulative flow diagram, a say/do ratio that keeps missing. Teams collect this data constantly and almost never sit with it.
Pull up throughput, the CFD, and the last two or three sprints of commitment-versus-completion. Walk each chart together, asking “what do we notice?” before “what does it mean?” Resist the urge to explain the numbers away in the first five minutes.
- What patterns show up in throughput over the last few sprints?
- What is our actual capacity, based on this data and not on what we wish it were?
- What external events moved these numbers that we should account for?
- Which single adjustment would do the most for predictability?
Leave with a commitment adjustment, a WIP limit, or a named capacity reserve for reactive work — expressed as a number, not an intention.
The One Where We Implement Our Technical Investment System
Run it when technical debt keeps coming up and keeps losing. The conversation stalls because the team has no mechanism for surfacing, sizing, and prioritizing technical work against feature work — so it gets argued case by case, and the case with a customer attached wins every time.
This retro builds the mechanism. Walk your technical investment approach section by section and agree on how your team will actually operate it. The output is process, not a backlog.
- How do we surface and capture technical investment opportunities?
- What criteria prioritize technical work against feature work?
- How much capacity do we reserve for it each sprint?
- Who owns the technical investment backlog?
- How do we show the impact of the investments we make?
Leave with documented mechanics, the working agreements that govern them, and the setup work — tooling, backlog grooming — scheduled across the next sprint or two.
The One Where We Make Working Agreements Explicit
Run it when the same friction keeps recurring and nobody can point to a rule it violates. Every team runs on norms; the question is whether they are written down or discovered mid-argument.
Brainstorm three to five candidate working agreements, pick exactly one, write it in language specific enough to be falsifiable, and schedule the review two sprints out. One agreement at a time. A team that adopts five agreements in an afternoon has adopted zero.
At the review, ask the questions that matter more than the original selection: Did we hold to it, and how often? Adjust, adopt permanently, or retire? Is there a sharper version? Did something more important surface instead? Can we make the tooling enforce it so the agreement stops needing willpower?
Leave with one agreement, a review date, and — at the review — an explicit disposition. “We forgot about it” is a disposition too, and a useful one.
Good agreements sound like this: “We don’t start new work with more than three items in code review.” “PRs get a first review within four business hours.” “We pair on anything in progress more than two days.” Bad ones sound like “we communicate well.”
The Facilitation Arc
Esther Derby and Diana Larsen gave the practice its spine: set the stage → gather data → generate insights → decide what to do → close. Every format above rides this arc. Skip a stage and you can predict the failure — skip gathering data and you get opinions, skip deciding and you get therapy with a whiteboard.
One scenario runs through all five stages below. A payments team’s throughput has fallen from fourteen completed items to nine over three sprints. Their say/do ratio sits at sixty percent: they commit to ten things and finish six. Leadership has noticed. Everyone on the team has a theory, and no two theories match. They run the metrics retro.
1. Set the Stage
Get every person in the room to say something in the first five minutes. Silence at minute five is silence at minute fifty. State the focus, state the time box, and state what this hour is not for.
In the room: the facilitator opens with, “We’re here to read three charts and leave with one change. We are not here to explain why the numbers look the way they do.” Then a one-word check-in around the room. Two people say “defensive.” That is data, and it arrived before the charts did.
2. Gather Data
Facts before feelings, and both count as facts. Put the charts up alongside a timeline of what actually happened — releases, incidents, people out, priorities changed mid-sprint. Ask what people notice. Write it down without interpreting it.
In the room: the CFD shows the code-review band widening every sprint. Someone adds that two of the three sprints absorbed a production incident. Someone else notes that the team started nearly every committed item in the first two days of each sprint. Three observations, zero conclusions.
3. Generate Insights
This is the stage teams skip, and it’s the one that earns the meeting’s cost. Move from noticing to causation. Ask Deming’s question — by what method? — and keep the search pointed at the system rather than the person. If the insight is “Dave needs to review PRs faster,” you haven’t found it yet.
In the room: the widening review band, the day-one start on everything, and the carryover are one pattern, not three. The team starts all ten items, review queues up behind them, unfinished work rolls into the next sprint, and the next commitment is inflated by work that was already half-done. It isn’t a capacity problem. It’s a WIP problem wearing a capacity costume.
4. Decide What to Do
One or two changes. Each with an owner, a size, and a place on the board. A retro that produces seven action items has produced none.
In the room: a WIP limit of three items in review, and fifteen percent of capacity held back for reactive work. Both become cards on the kanban — sized, owned, and counted against the team’s WIP like everything else, because that is the only way anyone will see them again.
5. Close
Name the decision, name the owner, name the date you’ll look at it again. Then spend two minutes on the retro itself: was this worth the hour? Teams that never inspect their own retro run the same broken one for years.
In the room: “We’re looking at this same CFD in two sprints and deciding keep, adjust, or kill.” Someone asks whether they should tell their director about the incident load. Yes — which turns out to be the most valuable thing the hour produced.
Where Retrospectives Go to Die
The silent room
The team answers “what went well” with the sprint report and “what didn’t” with nothing. This is not a facilitation problem you can solve with a better icebreaker. Retros collect garbage data without psychological safety, because the honest answer usually implicates someone with more power than the person holding it. Fix the safety problem first; the retro is a symptom, not the disease.
The venting session
Complaints are real signal, and a team that vents is at least a team that talks. But an hour that ends in catharsis has spent an hour arriving where it started. The facilitator’s move is the follow-up question: “What would have to change for that to stop happening?” Venting names the pain. The next question turns it into a decision.
The action-item graveyard
The most common failure, and the most fixable. Actions get captured in a retro doc, the doc gets closed, and nobody opens it again until someone is prepping for the next retro. Apply the working agreement this collection keeps coming back to: if it is not on the kanban, it does not exist. A retro action living in a doc is invisible work competing against visible work, and it will lose every time. Make it a card. Give it an owner. Let it consume WIP.
The rerun
The same theme surfaces sprint after sprint — the environment is flaky, the dependency team never responds, requirements arrive half-formed — and the team logs another action item aimed at a cause it doesn’t control. After the second recurrence, stop treating it as a team problem. It is a systemic blocker, and the retro’s output changes shape: not another action item, but an escalation carrying evidence. Three sprints of data, the cost in cycle time, and a specific ask. Bring a leader in to see it firsthand — a Gemba walk into the team’s retro does more than a slide about it ever will.
A team that has run retros for a year and cannot name one thing that changed didn’t hold twenty-six retrospectives. It held twenty-six status meetings with snacks.
Resources
- Esther Derby and Diana Larsen, “Agile Retrospectives: Making Good Teams Great” (Pragmatic Bookshelf, 2006) — the source of the five-stage arc and still the best book on the practice
- Psychological Safety — the precondition; without it a retrospective collects agreeable silence
- Working Agreements — where retro decisions become explicit and testable
- Gemba Walks — the leader-side counterpart; observe the work instead of receiving a report about it
- Flow Metrics Guide — the charts the metrics retro reads
- WIP Limits — the decision that comes out of a metrics retro more often than any other
- By what method? — the question that separates an insight from a complaint
Nerdy