Seeing How Well a Team's Story Points Align from One to Eight
Teams sometimes wonder whether their story point scale is working the way they intend. Are two-point stories really about twice the size of one-point stories? Do fives and eights still mean something useful, or has the team drifted into using them loosely?
One way to check is to compare completed stories by point value and look at the actual effort range for each bucket.
Reading the Chart
Each vertical range shows the shortest and longest completed story for that point value. The red line shows the median effort.
For this company, one-point stories ranged from 6 to 36 hours, with a median of 21 hours. Two-point stories ranged from 31 to 73 hours, with a median of 52.
If the one-point stories were perfectly calibrated, we might expect two-point stories to have a median of 42 hours. They came in at 52. Or perhaps the two-point stories were the better anchor, in which case the one-point median might have been closer to 26.
Most likely, neither anchor is perfect. The useful observation is simply that the team is close, but not exact, at the low end of the scale.
The three-point stories line up almost exactly: three times the one-point median is 63, and the observed median is 64.
Five-point stories also look reasonable once you remember how story points work as buckets. A story that feels like a four usually goes into the five-point bucket. If the bucket contains a mix of fours and fives, the average story in that bucket is closer to 4.5 points.
Multiplying 4.5 by the one-point median of 21 gives 94.5 hours, which is close to the observed median of 100.
The eight-point stories are where the drift becomes noticeable. An eight-point bucket often holds stories that feel like sixes, sevens, or eights. If the average story in that bucket is around seven points, seven times 21 hours would suggest a median near 147 hours. The observed median is only 111.
That does not mean the team did anything wrong. But it does suggest that some stories called eights may have been closer to fives, or that the eight-point bucket was being used differently than the rest of the scale.
Use the Data to Calibrate, Not Police
A more formal version of this analysis would use linear regression and look at the r-squared value to see how well the values fit. You can do that in Excel. But for a team conversation, a simple chart like this is often enough.
This team appears well calibrated through about five points and somewhat less consistent by eight.
If you collect data like this, be careful about how you introduce it. You will be measuring actual effort spent on completed stories, so team members may feel extra pressure to finish within their estimates.
If that happens, they may start padding estimates, which defeats the whole purpose. Make it clear that the point is calibration and learning, not judging individuals or rewarding estimates that match hours.
Show the graph to the team and ask what they notice. In this example, the team might decide they put eights on stories that were really closer to fives. Or they might decide the eight-point bucket contained more sixes than eights. Either way, the conversation helps the team make its point scale more consistent.
Most teams that try this find they can become reasonably consistent through about eight points. With awareness and a little practice, many teams can calibrate across a 1-13 range.
Beyond that, the numbers should be used with caution, or reserved for rough long-term questions such as whether a project is likely to take a couple of months or closer to a year.










