Measurement & verification

Why we adjust the baseline, not the results

If you compare this year's gas bill to last year's and call the difference a saving, you have not measured a retrofit. You have measured a winter.

Every energy retrofit ends with the same awkward question. You spent the money, you did the work, and now somebody in finance wants to know whether it worked. The temptation is to answer the easy way: pull last year's utility bills, pull this year's, subtract.

The problem is that between those two years, the weather changed, the building's occupancy changed, a tenant moved a server room in, and the utility estimated three meter reads. Any one of those effects can be larger than the retrofit. In a mild Calgary winter, the subtraction method will hand you a 15% gas saving you did nothing to earn. In a harsh one, it will bury a genuine saving completely and make a successful project look like a failure.

The International Performance Measurement and Verification Protocol exists to stop this. Its central idea is deceptively simple and, once you see it, difficult to unsee.

You do not compare the reporting period to the baseline period. You adjust the baseline to the reporting period's conditions, and compare against that.

What that means in practice

You build a regression model of how the building consumed energy before the retrofit, as a function of whatever actually drives consumption. Then, at the end of the reporting period, you feed that old model the new period's weather. It tells you what the un-retrofitted building would have used under this year's conditions. The gap between that and the metered reality is the avoided energy.

The retrofit gets credit for what it did. The weather gets credit for what it did. Neither borrows from the other.

Two fuels, two different drivers

At the Calgary tower we assessed last year, the two fitted baseline models look nothing alike, and that difference is itself a finding.

Natural gas (GJ/month) = 2.22 × HDD + 16.9 × days
Electricity (GJ/month) = 29.31 × days

Gas behaves exactly as you would expect. It correlates strongly with heating degree-days against a 17 °C reference temperature, returning an R² of 0.96. The 16.9 GJ per day term is the weather-independent baseload — domestic hot water and standing losses that continue in July.

Electricity has no weather term at all. Not a weak one; none. The model that best describes this building's electricity consumption is a flat 29.31 GJ per day, every day, regardless of whether it is January or July.

For a 23-storey tower with two water-cooled chillers, that is a diagnosis in its own right. A healthy building's electricity should rise in cooling season. This one's does not, because the plant was running at essentially full load all year — both chillers, all pumps, no staging, around the clock. There was no cooling-season peak because there was no cooling-season change. The flatness of that line is the staging failure, drawn in utility data.

Real data is messier than the model

Below is the twelve-month gas baseline the model was fitted to. Look at December.

November reads 1,713 GJ. January reads 2,069 GJ. December, sitting between them in the coldest part of a Calgary winter, reads 450 GJ — roughly a quarter of its neighbours. No building does that. What happened is almost certainly an estimated or missed meter read, with the shortfall recovered in the following billing periods, which is also why January and February both look slightly high.

This is the ordinary condition of utility data, and it is precisely why the baseline is a fitted model rather than a lookup table. A regression against degree-days absorbs a single bad month without collapsing. A month-by-month comparison would propagate that error straight into the savings number and nobody would notice.

It is also why we say what the reference temperature is, what the fit statistics are, and which months went into the fit. Someone should be able to disagree with us using the same data.

The number beside the number

Every model has an error band, and publishing it is not a hedge. It is the thing that makes the savings figure interpretable.

For this building: expected accuracy of ±9.7% on gas and ±13.6% on electricity, both at a 68% confidence level. Gas CV(RMSE) came in at 10.24% against a 25% limit; electricity at 13.6% against the same limit; normalized mean bias error on electricity was effectively zero.

Now put those two facts side by side. If a project claims a 5% electricity saving on a model with ±13.6% expected accuracy, that claim is indistinguishable from noise. It is not a small saving. It is not a saving at all, in any sense you could defend to an auditor.

This is the reasoning behind Option C's ten percent floor. Whole-building utility analysis only works when the effect you are looking for is bigger than the measurement error of the method. Go below that and you must isolate the equipment and meter it directly under Option B, or build and calibrate a simulation model under Option D. Both cost considerably more.

Watch it happen

The model below is loaded with the real fitted gas coefficients. Change the reporting-period winter and watch the adjusted baseline move with it — the savings figure stays honest because the comparison line moved, not the result. Then drag the savings slider under 10% and watch the method itself stop being defensible.

The third slider is worth understanding too. Non-routine adjustments cover everything the weather model cannot see: a floor going vacant, a tenant installing a server room, a change in operating hours. These are applied by hand, documented, and agreed with the client in advance of the reporting period rather than negotiated afterwards when the number comes in disappointing.

Agreeing the adjustment rules before you know the answer is the entire discipline, and it is the part most likely to be skipped.

Choosing the boundary

Option C measures at the utility meter, which means the boundary is the whole facility and everything inside it is in scope — including things that have nothing to do with your retrofit.

At this building, retail tenants are sub-metered for electricity, so their consumption can be handled explicitly. Retail natural gas metering was excluded from the boundary because it was not accurate enough to rely on. That is a judgement call, it is documented in the plan, and it is the kind of decision that determines whether a verification survives review.

Electricity and gas were also fitted over different twelve-month windows — gas from November 2022, electricity from July 2024. That looks like sloppiness and is not: ASHRAE Guideline 14 permits each fuel to be fitted over the period that best characterizes its own governing driver. Forcing both onto a common window would have degraded one of the two models for the sake of tidiness.

Why any of this matters commercially

Verified savings are what incentive programmes pay against, what performance contracts settle on, and what carbon disclosure frameworks will accept. An unverified saving is a marketing claim.

But the more practical reason is internal. A building that has a fitted baseline model has a permanent early warning system. When consumption drifts above the adjusted baseline six months after a recommissioning project, you find out — and you find out that a setpoint was reverted, or a schedule was overridden during a service call and never restored. Without the model, that regression is invisible until somebody happens to look at a bill.

Most savings from controls work are lost this way. Not because the work was wrong, but because nothing was watching.

See the full plan

The verification plan described here covers eight energy efficiency measures totalling 5,522 GJ per year — about 25% of total baseline consumption — at 444 5th Avenue SW in Calgary.