Saturday, October 23, 2010

Corsi Corrected for Schedule Difficulty

While another year of hockey is finally underway, the 2010-11 season is still very much in its infancy. The schedule has yet to reach the 100 game mark, with no team having played more than a handful of games. This being the case, drawing conclusions on the basis of the results thus far can be difficult. The sample size with which we have to work just isn’t large enough.

To illustrate this, consider the league standings as of Friday, October 22nd, following the completion of the 97th game. The range in standing points is 7, with a standard deviation of 2.08. In assigning each team the same winning percentage, setting home advantage at 5%, giving each game a 22% chance of going past regulation, and simulating the first 97 games of the schedule 1000 times under these conditions, the following is obtained.

In other words, virtually all of the team-to-team variation in standings points at this stage of the year is the product of randomness. (For interest’s sake, only about half of the variation in standings points over the course of an entire season can be accounted for by luck.)

As shots in hockey are relatively frequent events, it makes much more sense to rely on a shots-based metric in order to get a sense of how each team has performed thus far. But which metric in particular ought to be used? And which adjustments, if any, are necessary?

Corsi – which includes all attempted shots and therefore better attenuates any sample size concerns – serves as the best fit for this exercise, as opposed to either Fenwick or shot ratio proper. However, two adjustments are necessary. Firstly, playing to the score effects – which are known to bias the shot clock in favour of the trailing team – ought to be controlled for as much as possible. This is especially true this early in the season, as some teams will have played with the lead for much longer periods than others. Ideally, one would restrict the sample to shots attempted at even strength with the score tied in order to get around this problem. However, because of the sample size concern identified above, using Corsi with the “score close” – defined as whenever the score is within one goal in the first or second period, or tied in the third period or overtime – is to be preferred.

Secondly, a team’s Corsi depends not only on its own ability to outshoot the opposition at even strength, but also on the ability of its opponent in this respect. At this point in the year, few teams have played what could be reasonably described as a balanced schedule. Thus, regard should be had to the fact that some teams have faced stronger or weaker opponents with respect to Corsi through incorporating some sort of correction for strength of schedule.

The table below shows each team’s Corsi percentage with the score close and how that changes once an adjustment for strength of schedule is applied.

It’s important to note that, at this point in the year, roughly 43% the team-to-team variation in Corsi percentage with the score tied (raw, not adjusted) can be attributed to luck. Accordingly, some teams will see their ranking change significantly between now and the season’s end. If forced to predict, I’d wager that, relative to underlying talent, the Devils, Bruins, Capitals, Blackhawks and Sharks are better than these rankings suggest. Conversely, I’d wager that the Avalanche, Canadiens, Rangers, Panthers and Flyers are worse.

Friday, May 28, 2010

Stanley Cup Final Prediction and Probabilities


[For an explanation of the table and how the odds were computed, see here].

Not much to say here.

If the regular season was any indication, all three of Chicago's playoff opponents were better teams than the Flyers.

If that doesn't reflect the imbalance between the two conferences, I'm not sure what does.

CHI in 5.













Saturday, May 15, 2010

3rd Round Playoff Predictions and Probabilities


[For an explanation of the table and how the odds were computed, see here].

S.J - CHI

Although Chicago was the better team during the regular season, the Sharks have, to my eye, looked more impressive through the first two rounds. Chicago has simply not been anywhere near as dominant as I would have anticipated.

That said, I thought Chicago was the best team in the conference before the playoffs started, and, while their recent play has produced some doubt in that regard, that belief still holds true.

Thus, I'm going with the Blackhawks to win the de-facto cup final.

CHI in 6.

PHI - MTL

While both of these teams have received some good fortune in order to be where they are right now, the Habs playoff run has been more luck driven. Given that luck doesn't persist over time, this works in the Flyers favor.

Philadelphia may be the weakest opponent that Montreal has faced thus far, but they're still the better team.

I've enjoyed Montreal's playoff run immensely. During my tenure as a serious fan, I had never, before this year, had the opportunity to watch my team advance beyond the second round. Although the circumstances of their advancement leave much to be desired - as any self-respecting fan would prefer to see his team win on merit -, I'm glad that it's finally happened.

I have a feeling that it ends here, though.

PHI in 6.





Wednesday, April 28, 2010

2nd Round Playoff Predictions and Probabilities




[For an explanation of the table and how the odds were computed, see here].

S.J - DET


This series is interesting in the sense that there is no obvious favorite. The Sharks had the better regular season goal ratio by a fair amount, and have about a 65% chance to win if the odds are computed on that basis. On the other hand, the Wings had the better underlying numbers. While both teams were very good at generating shots on the powerplay and moderately good at shot prevention on the penalty kill, the Wings were better at outshooting at EV with the score tied.

I've included an excel document below that contains a list of series from 1993-94 onward where one of the teams had the better pythagorean expectation, and the other the better shot ratio. The team with the better pythagorean expectation is listed under the column heading 'T1', whereas the team with the better shot ratio is listed under the column heading 'T2'. 'W%' denotes pythagorean expectation, whereas 'SR' stands for shot ratio. 'Result' indicates which team won the series. 'W' indicates that the team with the better pythagorean expectation won, while 'L' indicates that the team with the better shot ratio won.

(I realize that shot ratio and the underlying numbers are distinct metrics; however, the data necessary to compute an expected winning percentage based on the underlying numbers just isn't available for the seasons in question. In lieu of that, I think that shot ratio provides an adequate proxy.)



Overall, there were 65 series that satisfied the above criteria. Of those series, the team with the better pythagorean expectation won 35 times, while the team with the better shot ratio won 30 times. This bodes well for the Sharks, I think.

The average difference in shot ratio between the two teams was about 0.12, which is almost identical to the difference between the Sharks and Wings. The average difference in pythagorean expectation was about 0.04, which is less than the 0.07 separating the two teams. Again, I think that this works in San Jose's favor.

On the other hand, the Wings almost certainly aren't a true talent 0.53 team, and it would be foolish to regard them as such.

All things considered, I think that this matchup is pretty close to a cointoss. I'm going with the Wings, if only because I think that the underlying numbers method provides a better measure of a team's true ability than does pythagorean expectation, even though there may or may not be an empirical basis for that viewpoint.

DET in 7.

CHI –VAN

The Canucks are a good team, but I can't help but get the sense that they're a tad overrated. I was browsing Hfboards the other day and I noticed that some 60% of the posters there have picked Vancouver to win the series. To be sure, some of that has to do with the fact that Canucks fans outnumber Hawks fans among HF users. Even so, I found the poll results interesting as the numbers suggest that Chicago is the better team. As posted above, the Hawks are about a 60% shot to win on the basis of adjusted winning percentage, and about a 70% shot if the underlying numbers are used.

The two teams were actually pretty close to one another in terms of regular season goal differential, but the Hawks were much, much better at outshooting. Chicago led the league with a shot ratio of 1.36 (awesome), whereas the Canucks were tenth at 1.05 (meh).

What interests me is how often a playoff team in the Hawks position has performed historically in terms of series wins and losses. That is to say, if two teams are facing one another in the playoffs, and one team has the better regular season shot ratio by a large margin (say, at least 0.2 better), but is only slightly better in terms of pythagorean expectation (say, no larger than 0.08), how often does that team end up winning?

Looking strictly at playoff results between 1993-94 and 2008-09, I found 32 series that met these criteria. I've arranged the series according to date in the excel document below. The headings may require some explanation. 'T1' denotes the team with the better shot and goal ratio, whereas 'T2' denotes their opponent. W% stands for adjusted winning percentage, and SR stands for shot ratio. The 'Results' column indicates which team won the series. 'W' indicates that the team with the better goal and shot ratio won the series, whereas 'L' indicates that the other team won. The bottom column shows the average adjusted winning percentage and shot ratio for the T1 and T2 teams, respectively. As it turns out, the T1 and T2 teams differed, on average, by about 0.03 in adjusted winning percentage and by about 0.3 in shot ratio, which, in both cases, is virtually identical to the gap separating the Hawks and Canucks.



All in all, the T1 team won 19 out of the 32 series, or 58%. That's hardly overwhelming and, to be honest, I would have expected that number to be higher. If the historical results are to given any weight at all, Chicago's chance of winning the series is probably closer to 60% rather than the 70% figure generated by the underlying numbers model.

In any event, the historical results are consistent with my general point that the Hawks ought to be the favorite here. The Canucks have a reasonable chance to win, but it's not somewhat that should be expected in the sense of being more likely than not.

CHI in 6.

PIT-MTL

This pick doesn't require too much deliberation. The Pens might be the best team in the conference, whereas the Habs are easily the weakest squad to advance. Pittsburgh is the heavy favorite regardless of whether the odds are determined through each team's pythagorean expectation or through the underlying numbers. That said, 29% ain't trivial and, as we observed last round, anything can happen over the course of a best-of-seven series.

PIT in 5.

BOS-PHI

I think that these two teams are relatively equal, but that the Bruins are slightly better. It's hard to pick against a team that's as good territorially at even strength as Boston is, even for a Habs fan such as yours truly. To add to that, Savard is expected to return for the series, and that should help them. I expect them to advance.

BOS in 6.

Tuesday, April 27, 2010

The Repeatability of Special Teams Performance

In my post on playoff probabilities, one of the methods in which I calculated each team's expected winning percentage was on the basis of the underlying numbers.

Under this model, shot volume on the powerplay, shot prevention on the penalty kill, as well as penalty differential, were incorporated as determinants of special teams goal differential. However, neither shooting percentage on the powerplay nor save percentage on the penalty kill were used as predictors.

Initially, my intention was to include both variables within the model. However, after looking at the relationship between team powerplay shooting percentage in even numbered games and team powerplay shooting percentage in odd numbered games in the 09-10 regular season, I discovered that there was essentially no correlation. I then did the same thing for the 07-08 season, and the result was the same: no relationship.

I found this to be unusual, given that I had looked at the distribution of powerplay shooting percentage in the past and found that the team-to-team spread was somewhat broader than what one would expect if there was no skill component. Nevertheless, my exercise had revealed the absence of any split-half correlation, thus necessitating the exclusion of PP S% from the model.

(As mentioned above, I also excluded PK save percentage, even though I had not specifically examined its repeatability. This was somewhat unjustified given that, as discussed below, team PK SV% is somewhat repeatable. However, the regression is fairly strong and, even though I ought to have taken it into account, it's exclusion didn't affect things too greatly.)

In any event, my curious findings prompted the following question: To what degree is special teams performance repeatable?

Real Effects

Vic Ferrari had an excellent post about a year ago where he looked at the various components of team even strength performance -- specifically, shooting percentage, save percentage, and shot differential -- and determined the extent to which each component was repeatable. Specifically, his method involved looking at each team's shooting percentage, save percentage, and shot differential, all at EV with the score tied, in 38 randomly selected games from the 2008-09 season. He then looked the same variables over a separate 38 game sample, and determined the correlation between the two sets of games. The exercise was then repeated over 1000 simulations.

The rationale behind the exercise is a simple one -- as expressed by Vic, "if an element of nature is affected by something other than randomness, that it should sustain itself from one independent sample to another." Thus, if the observed correlation is significantly non-zero, it can be assumed that the variable is at least partly determined by factors other than luck. On the other hand, if the observed correlation is insignificantly different than zero, then fluctuations in the variable are assumed to be primarily luck driven.

I decided to apply a similar technique in order to determine the degree to which the components of special teams performance are governed by 'real effects.' Specifically, my methodology involved the following:
  • I obtained special teams data at the team level for each season from 2003-04 to 2009-10
  • Within each season, I looked at team performance on specials teams at the level of individual games
  • In particular, I looked at the following variables: powerplay shooting percentage, penalty kill save percentage, powerplay shot rate (shots for divided by time on ice), penalty kill shot rate (shots against divided by time on ice), and powerplay ratio (the ratio of powerplays drawn to powerplays conceded)
  • However, shooting rates were not examined for 2007-08, 2008-09, and 2009-10, as I was not able to obtain data on PP TOI and PK TOI for those seasons
  • Empty net goals were excluded when calculating shots and goals
  • For each team, I randomly selected 20 home games and 20 road games, combined the two sets of games, and looked at how that team performed within that sample with respect to the above stated variables
  • I then did the same thing for 40 other randomly selected games (again, consisting of 20 homes and 20 road games)
  • I then looked at the correlation between the two sets of games for each of the listed variables
  • I repeated the exercise 1000 times, for each of the six seasons
The results


I should note that the final highlighted column shows the averaged value for each variable.

As indicated by the table, both generating shots on the powerplay and preventing shots on the penalty kill appear to be largely ability driven measures. The same applies to drawing more powerplays than the opposition.

Not surprisingly, both PP S% and PK SV% are less ability driven than the other three variables. It's worth noting that PK SV% appears to be more reliable than PP S%. I presume that this can be attributed to the influence of the goaltender on PK SV%.

Wednesday, April 14, 2010

Corrected Playoff Probabilities

Because I forgot to include EV goals when calculating each team's corsi with the score tied, I decided to re-run the UNDERLYING #'s simulation using the corrected probabilities.

The results aren't too different.

Expected Winning Percentage by Team

In response to a question raised in the comments to my post on playoff probabilities, I figured that it would be useful if I posted each team's expected winning percentage according to the two described methods.

The teams are ranked according to pythagorean winning percentage.


EDIT:

I've altered the chart so as to include even strength goals in the calculation of Corsi with the score tied. The values don't change all that much -- in fact, hardly at all, but I figured that I'd post it if only for accuracy's sake.