Volatile velocity, wildly different completed-work totals from one sprint to the next, undermines the forecasting value velocity is supposed to provide. Before reaching for stabilization tactics, it's worth understanding that velocity swings are usually a symptom, and treating the symptom without addressing the underlying cause rarely produces lasting stability.
The usual causes, in rough order of frequency
Inconsistent estimation is the most common cause, and it's rarely about the team being bad at estimating in any absolute sense: it's about estimation criteria drifting over time, or being applied inconsistently across different types of work, or different team members implicitly using different mental scales for the same numbers. A "5" estimated by one person and a "5" estimated by another aren't necessarily comparable if they've never explicitly calibrated against each other.
Poorly sized backlog items are the second most common cause. Items that are too large carry more estimation uncertainty, and when a handful of oversized items land in the same sprint, velocity swings on the outcome of just those few items rather than reflecting the team's actual typical output. Consistently splitting work into smaller items reduces the variance contributed by any single item's estimate being wrong.
Unaccounted-for interruptions (production incidents, urgent unplanned requests, context-switching onto other work) eat into capacity in ways that don't show up in the plan but absolutely show up in the result. Sprints that look identical on paper can have very different actual available capacity depending on how much unplanned work landed.
Team composition changes (someone on leave, someone new who isn't yet at full productivity, a team member pulled onto another initiative for part of the sprint) directly affect available capacity in ways that a velocity number calculated from headline team size doesn't capture.
What actually stabilizes velocity
Calibrate estimation explicitly and periodically, not just once during initial team formation. A short, recurring practice of estimating a few historical items together and discussing disagreements keeps the team's shared scale from silently drifting apart over time.
Track and account for unplanned work separately, rather than letting it silently eat into planned capacity without being visible anywhere. Teams that measure the ratio of planned to unplanned work over several sprints can plan future sprints with a realistic buffer, instead of being surprised by the same pattern every time.
Aim for a consistent range of item sizes, actively splitting anything unusually large before it enters a sprint, rather than accepting outsized items and hoping the estimate holds.
Look at velocity over a rolling average of several sprints, not sprint-to-sprint. Some volatility is inherent and not a sign of a problem: a three- or five-sprint rolling average smooths out normal noise and reveals genuine trends, whereas single-sprint comparisons often just amplify statistical noise into an apparent crisis.
Stable velocity isn't really the goal in itself. It's a byproduct of a team with consistent estimation habits, reasonably sized work, and visibility into what's actually eating its capacity. Chasing the number directly, without addressing what's driving its instability, tends not to work.