Wednesday, June 27, 2012

A new homogenised daily temperature data set for Australia




The ACORN-SAT station at Butlers Gorge in central Tasmania
 
A new homogenised daily temperature set, the AustralianClimate Observations Reference Network – Surface Air Temperature (ACORN-SAT)data set, has recently been released by the Australian Bureau of Meteorology. This data set contains daily data for 112 stations, with 60 of them extending for the full period from 1910 to 2011, and the others opening progressively up until the 1970s. (1910 is taken as the starting point, as it was only with the formation of the Bureau of Meteorology in 1908 as a federal organisation that instrument shelters became standardised). 

The new data set applies differential adjustments to different parts of the daily temperature frequency distribution, using an algorithm which matches percentile points in the frequency distribution with those at reference stations before and after an inhomogeneity. This takes into account the fact that some inhomogeneities in temperature records have different impacts in different weather conditions – for example, if a site moves from a coastal location to one further inland, the difference in overnight minimum temperature will normally be greater on clear, calm nights than on cloudy, windy nights, and clear, calm nights are also likely to be the coldest nights (at least in the mid-latitudes). If such effects are not accounted for, adjustments which homogenise mean temperatures may not homogenise the extremes. 

The method is conceptually similar to that of Della-Marta and Wanner (2006), while differential adjustment methods of this type have also been developed by a number of other authors. The ACORN-SAT data set, however, is believed to be the first implementation of such methods in a national, year-round data set. 

A detailed evaluation of the method was carried out before finalising its implementation. One of the interesting findings from the evaluation was that, while some earlier studies had suggested that reference stations required a very high correlation (0.8 or above) for use in daily adjustment methods, the ACORN-SAT evaluation found that correlations of 0.6 or above produced satisfactory results. The reasons for this difference are yet to be fully evaluated. Two potential explanations are that ACORN-SAT uses multiple reference stations (normally 10) whilst the other methods were evaluated using a single reference station, and that the ACORN-SAT evaluation took place using real data, whilst the other evaluations used artificial data sets whose properties may not necessarily match real-world inhomogeneities. This result is critical for the use of such methods in Australia, where observing networks are sparse in many areas by the standards of developed countries; had a minimum correlation of 0.8 been necessary many locations would have had no available reference stations.
Documentation of the data set at all points was considered to be critical. The homogenised and raw daily data, all of the transfer functions used in the adjustments, and other relevant documentation, are available on the ACORN-SAT website. This information would allow the data set to be reproduced from the raw data should anyone wish to do so.  

The methods used in the development of the data set have recently been published in a paper in the International Journal of Climatology. A more extensive technical report has also been published by the Centre for Australian Weather Climate and Research and is available on the ACORN-SATwebsite. Other material published on the website include a second technical report which compares outcomes from the new data set with other national and international data sets, and a station catalogue containing metadata for the 112 locations. 

(One of the author’s personal goals is to visit all 112 locations; the 91st, Cape Moreton in Queensland, was visited in April. Occasionally his ambitions have exceeded his vehicle’s capabilities – drowning his last car while attempting to get out of a remote site in the far north of Western Australia).

The new data set will be used for operational climate analysis by the Australian Bureau of Meteorology, including reporting of annual national temperature anomalies. (Preliminary analyses indicate that the warming trend in the new data set, about 0.9°C over the 1910-2011 period, is similar to that in the data set previously in use). It will also allow, for the first time, analyses of century-scale changes in Australian temperature extremes. Such analyses are expected to be released over the next few months. 

References

Della-Marta, P.M. and Wanner, H. 2006. A method of homogenizing the extremes and mean of
daily temperature measurements. J. Clim., 19, 4179-4197.

Trewin, B.C. 2012. A daily homogenised temperature data set for Australia. Int. J. Climatology, published online 13 June 2012.

Trewin, B.C. 2012. Techniques used in the development of the ACORN-SAT dataset. CAWCR Technical Report 49, Centre for Australian Weather and Climate Research, Melbourne.

Thursday, April 12, 2012

Request for help: Identifying, prioritizing and digitizing International hard copy holdings held at NOAA NCDC

Update 4/26: The googledocs share spreadsheet linked below has been updated and simplified which will hopefully make this easier for people to engage in. There are also plans to host these images online soon and hopefully allow anyone interested to digitize and submit the digitized records for inclusion in the data holdings.

Colleagues at NCDC have recently embarked on a project of truly epic proportions. To inventory, image as necessary, and eventually digitize (to the extent useful unique information exists in them) the large volume of international holdings (>2000 boxes) held in hard copy in the NCDC basement.  That is a lot (an awful lot) of boxes ...
These consist of a huge range of different, primarily land in-situ, meteorological holdings. Some may be unique, others may exist elsewhere already as images or have been digitized already. Below are just a couple of teaser images ...
We would value yours and others' collective help in prioritizing the imaging and digitization of these holdings, telling us what has already been done and in actually doing some of the work. 

So, with that ...

An editable form of the current version of the spreadsheet summary with about 15% inventoried (Africa and S. America largely) and 1% imaged (highlighted yellow) is available at:


Please direct edit this, I do not want to have to manage 50 versions of the same spreadsheet or merge them ...! Alternatively you can leave a comment below if you are more comfortable doing that.

Edits can request further forensics (which stations, when, what), give reasons for interest in the data, offer to digitize images from that set of data etc. etc.

In terms of next steps, in coming weeks the land data images taken will very likely start to be hosted on the International Surface Temperature Initiative databank ftp site at NCDC in appropriate stage 0 (raw data imagery) directories. Then, ideally, we will get some help in digitizing these at which point we can start to make the digital data available without restriction through that databank portal and pull it through to NCDC's products as well as allowing others to use and investigate it. This will potentially help us fill in significant gaps in our knowledge of climate change in many regions and periods.

If there is significant interest I will update the spreadsheet with new inventories of boxes periodically.

Some resources which might help in this task should you wish to partake in it:

http://docs.lib.noaa.gov/rescue/data_rescue_home.html - current NOAA foreign data library imagery
ftp://ftp.ncdc.noaa.gov/pub/data/globaldatabank/ - ever growing resource of digital data for land stations - feel free to play ...
ftp://ftp.ncdc.noaa.gov/pub/data/globaldatabank/monthly/stage2/INVENTORY_ALL_monthly_stage2  - list of current stations and periods of record (incl. lots of duplicates)
More on the databank effort, including submission guidance for digital holdings that may not already be there, can be found at http://www.surfacetemperatures.org/databank . We are close to releasing a first version of the databank, but it is not too late to receive data submissions for consideration in the first version ...

Many thanks in advance for any help received in this task.

Wednesday, February 1, 2012

Initiative progress report posted

The first Initiative progress report has been posted here. This details the acheivements, progress against agreed workplans, and any issues. They will be produced annually. Although primarily for reporting purposes to 'sponsors', constructive feedback from others is welcome on this blog. We will do our utmost to take any comments on board in future activities.

Tuesday, January 10, 2012

More on benchmarking - COST HOME paper released

Continuing on from the theme of the last post the main paper of the COST Action HOME has been published today in Climate of the Past. This predominantly European effort (COST is a European cooperation mechanism) looked to assess a very large suite of approaches to homogenization efforts using a number of test cases which consisted of either real data or synthetic data which had known issues to the data creators for relatively small (compared to the size of the global data holdings) case study regions. Both surface temperature and precipitation were considered. Very many candidate approaches to homogenization were considered and assessed in a consistent manner allowing robust conclusions about the relative merits and strengths and weaknesses of different approaches.

The challenge for the surface temperature initiative is to build upon this by creating global rather than smaller regional benchmarks and to have a  similar number of algorithms applied to the much larger global database. Only through looking at the problem in multiple different ways and consistently benchmarking these approaches we will be able to properly (adequately, robustly - take your pick) assess the true uncertainties in our knowledge of global and regional climate change and variability.

For more in-depth information on this study please see the lead author Victor Venema's blog post at http://variable-variability.blogspot.com/2012/01/new-article-benchmarking-homogenization.html and links therefrom.

Friday, January 6, 2012

Benchmarking and Assessment applied to USHCN

Firstly, an up-front caveat, I am third (last) author on this paper.

Today the Journal of Geophysical Research has published a paper that applies the benchmarking and assessment principles of the surface temperature initiative to the USHCN dataset of land surface air temperatures.
Williams, C. N., Jr., M. J. Menne, and P. Thorne
Benchmarking the performance of pairwise homogenization of surface temperatures in the United States
J. Geophys. Res., VOL. 117, D05116, 16 PP. doi:10.1029/2011JD016761, .  (behind a paywall - sorry) (URL and details updated 10/31)

Update 1/19/12: NCDC have provided a version commensurate with AGU copyright guidance (and copyright AGU) at ftp://ftp1.ncdc.noaa.gov/pub/data/ushcn/v2/monthly/algorithm-uncertainty/williams-menne-thorne-2012.pdf (updated 10/31/12 following ftp server move). The paper is also highlighted on their "what's new" page at http://www.ncdc.noaa.gov/oa/about/temperature-algorithm.pdf .

The analysis takes the pairwise homogenization algorithm used to create the GHCN and USHCN products and does two things.

Firstly, it identifies a large number of decision points within the algorithm that do not have an absolute basis and allows these to vary. A good climate science analogy here is the perturbed climate model runs of the climateprediction.net project and other similar projects. These decision points were varied by random seeding of values to create a 100 member ensemble of solutions. This at least starts to explore the parametric uncertainties within the algorithm (and any interdependency's) and their implications for our understanding of the observed temperature record evolution. However, what it does not necessarily do is give us any better an idea as to what the true climate evolution may have been. Which brings us on to the second innovation and the focus of this post ...

Secondly, in addition to running on the observations these ensembles were run on a set of eight analogs to the USHCN network. These consisted of sets of data which directly mimicked the observational availability of the USHCN network itself through time. They were based upon climate model runs from a range of models and a range of forcing scenarios. This ensures some 'plausible' spatio-temporal coherency to the large-scale temperature fields. On top of these were super-imposed additional differences to mimic potential random and systematic influences of instrumental and operational artifacts. A set of distinct possibilities were explored across the eight worlds ranging from no systematic biases at all (highly improbable but a useful 'algorithm does no harm' test) through to a scenario where the network was bedevilled with very many largely small breaks with a sign-bias tendency - a situation which any algorithm would find hard to cope with. Unlike the real-world these analogs afford a luxury of knowing the true answer so that it is possible to actually benchmark and understand the fundamental algorithm performance and any limitations. Then it is possible to re-evaluate the real-world results afresh with these new insights gleaned from such realistic test-beds.

The results from the analogs were broadly encouraging. First and foremost when applied to the data with no breaks added virtually no adjustments were made and the impact on large-scale averaged timeseries and trends was so minuscule a magnifying glass would be required to tell the difference. So, in the implausible eventuality that the raw data are bias free the algorithm really would do no harm. For the other analogs the performance was mixed. Easier cases where breaks were bigger and metadata better it did better. Harder cases it fared worse. Where there was no overall sign bias in the applied breaks the ensemble was spread relatively evenly around the starting data. But where there was an overall bias in the raw data presented to the algorithm it consistently moved the overall data in the right direction but rarely far enough. The implication being that this uncertainty is effectively one-tailed. The chances of over-shooting the adjustments is substantially smaller than the chances of under-shooting the required adjustments.

So, what does the real-world look like when reassessed through this new understanding?

For minimum temperatures over the longest timescales the trends are spread around the raw data. But in the periods 1951 onwards and 1979 onwards it is spread distinctly either side of the raw data. Post 1951 trends in the raw data exhibit too little warming, and post-1979 too much. This is consistent with prior understanding of the biases of the change of time of observation (largely 1950s-1970s; spurious cooling) and move to MMTS sensors (early to mid-1980s and often associated with a microclimate relocation; spurious warming) on Tmin.

For maximum temperatures the ensemble consistently precludes the raw observations over all considered timescales. The raw data are almost certainly biased and show too little warming. Over all periods the ensemble of solutions show more warming than the raw data. Again, this is consistent with current understanding of the impact of time of observation biases (again a spurious cooling) and the transition to MMTS which unlike for minimum temperatures imparts a spurious cooling effect. The less than encouraging implication from the analogs, however, is that in the operational USHCN algorithm we are substantially more likely to be under-estimating the required adjustments, and hence rate of warming in maximum temperatures, than over-estimating it.

Is this the last word on the issue? Certainly not. As this was a first step along this path the analogs were perhaps not as sophisticated as would ideally be the case. And here we hope the benchmarking and assessment group can provide more realistic (and global) analogs later this year. Further, this considered solely parametric uncertainty - varying choices within one algorithm. The larger and more difficult uncertainty to understand is the structural uncertainty that would result from applying fundamentally distinct approaches to the same problem and allow a better exploration of the possible solution space. And here we have to look to you, readers, to develop new and novel approaches to homogenizing the data and submit them to the same raw data and analogs to allow consistent benchmarking and better understanding.

Finally, over coming weeks we will be hosting code, data, and metadata (including the analogs) online. I'll provide an update when it is all up there but given that its several Gb and requires fitting around other duties its not going to be instantaneous by any stretch.

Comments are welcome, but please remember that this is a strictly moderated blog and to follow the house rules.

Thursday, December 15, 2011

Initiative overview paper published in the Bulletin of the American Meteorological Society

The first peer-reviewed paper describing the end to end envisaged International Surface Temperature Initiative scope has been published this week by the Bulletin of the American Meteorological Society. It is an Open Access paper available at http://journals.ametsoc.org/doi/pdf/10.1175/2011BAMS3124.1.  Any comments, queries or offers of effort are very welcome either through this blog or the general enquiries email general.enquiries@surfacetemperatures.org. Updates since the paper submission can be found at www,surfacetemperatures.org.

Thursday, November 10, 2011

GHCN-M v3.1.0 – showing the value of engaging with software engineers


NCDC have just released version 3.1.0 of the GHCN product, detailed in a tech note, as documented in the dataset paper of their global Land Surface Air Temperature product – the Global Historical  Climatology Network Monthly. This release does two things.

Firstly it incorporates an array processing algorithm that significantly speeds up the processing which will enable NCDC to process the much larger databank holdings upon its first version release in early-to-mid 2012 to form a yet more comprehensive estimate of the global Land Surface Air Temperature evolution.

Secondly, and the focus of this post, is that it incorporates a set of five process bug fixes, four of which were discovered in the homogenization algorithm as a result of an effort undertaken by Daniel Rothenberg sponsored by the Google Summer of Code and mentored by the Climate Code Foundation. The final bug was discovered as a result of carefully checking for similarly based bugs which essentially related to array compression / non-compression for missing values on passing between routines. That bugs exist in what is several thousand lines of code is hardly surprising. In fact it would have been far more surprising if it had been discovered that there were no bugs. Daniel visited NCDC as part of his project and the bugs were discussed at length with relevant NCDC staff and fixes have subsequently been undertaken, extensively validated, and their impacts on the analysis documented.

The bottom line impact on the global mean trend is a difference of less than 0.002K/decade – below the typically quoted global mean estimate precision of 2 decimal places and two orders of magnitude less than the reported centennial scale global-mean Land Surface Air Temperature warming rate from this dataset. Equally global annual means show negligible differences. Differences at the station level are almost always below 0.2K/decade with effectively zero mean change. So, whilst the bug fixes were important from both a science and process perspective they do not significantly alter our current understanding of changes in climate at the largest space and longest timescales.

What this does provide is an example of the very real potential value in openness and transparency, in code replication, and in working in positive partnership to resolve the issues that arise. Daniel aims to continue working on his port of the algorithm to python and it will be of great interest to see what other benefits may accrue.

NCDC have released the old (v.3.0.0) and new (v.3.1.0) versions of the homogenized data (in frozen form) and other relevant metadata (with ongoing additions) at ftp://ftp.ncdc.noaa.gov/pub/data/ghcn/v3/archives/.

Monday, October 31, 2011

WCRP OSC thoughts


WCRP OSC was a very large conference, mainly poster based. The sheer volume of posters was over-whelming. Plenary talks were generally very good, whereas the parallel sessions were a mixed bag with a number of real gems. Being talked at for ten hours a day is too long though and interest inevitably wanes. Wireless connections certainly aren't a help in that regard allowing people to attend without really truly being in attendance. I was presenting four posters across two sessions (one on my other 'hobby' - the GCOS Reference Upper Air Network) on the Tuesday morning and giving a talk on the surface temperature initiative the Tuesday afternoon.

I warmed up for this by asking a question in the c.2000 attendee observations plenary session first thing on Tuesday - nerve wracking in its own right. Two of the plenary speakers had bemoaned the lack of agreement between estimates for many variables and stated to be ‘scared’. I pointed out that this was an inevitable consequence of making measurements that were not traceable to measurement standards and that I was instead encouraged to see multiple estimates as this was the only way we could ascertain what could / could not be said. None of the speakers responded so either it was an awful point to make or they did not wish to respond.

I spent the majority of the poster time around the three surface temperature initiative posters. Like many of the posters they were in a corner but there was still reasonable interest and a number of potential data leads were identified. Roughly half of the 50 data request cover letters and data submission guidelines hardcopies were taken. Most of the discussants were supportive although inevitably some raised the Berkeley effort and whether this now obviated the need for the initiative as a whole. This gave an opportunity to clarify the holistic nature of the enterprise and how the Berkeley effort, if published(!), would simply constitute one important contributing component. It was stressed that science and society are interested in more than the global centennial timescale trend and that differences would be greater at smaller space and timescales. It was also stressed that consistent benchmarking was necessary to understand differences more robustly. See also Steve Easterbrook's take on the poster that he was presenting on benchmarking.

The afternoon talk was given in a parallel session with probably 300-500 people (it felt like the latter!) in attendance. It was a little bit rabbit in the headlights for the first half although better towards the end. There were at least three (maybe four) questions from the floor. There were then several people who had questions after the end of the session that kept me busy for the full half hour coffee break and beyond. These gave a chance for a much smaller audience to expand on various aspects – especially crowdsourcing. Questions regarding whether data holdings known to a given individual were already there highlighted the need to clarify that we wished to get hold of any and all data and that the databank processing will be designed to account for such redundancy in an open and transparent way. Regardless, a concatenated master-list of current holdings at stage 2 level was requested. This has now been added to the databank prototype.

Tuesday, October 18, 2011

ISTI at WCRP OSC

Or acronym soup?

For those who follow this effort and will be attending next week's WCRP OSC conference in Denver, Colorado, we will have a talk in Session B4 (Tuesday @15.00) and four posters in Session C13 (Tuesday morning). The posters are also posted at www.surfacetemperatures.org and the oral presentation will be uploaded there also after the event. Please do pop by and say hello.

Monday, October 17, 2011

ISTI Meeting Reports: 4th ACRE Workshop and GCOS Steering Committee meeting

I've just got back from presenting the work of the International Surface Temperature Initiative at both the 4th ACRE Workshop (Utrecht, Netherlands) and to the GCOS Steering Committee meeting (ECMWF, UK). The presentations and meeting reports are now hosted at: http://www.surfacetemperatures.org/background.

Kate

Wednesday, July 27, 2011

You may have noticed we have a logo now ...

Things are also starting to take shape in revamping the website. Any comments on this and suggestions as to how to make surfacetemperatures.org more useful gratefully appreciated.

Thursday, July 14, 2011

Overarching implementation plan published

We have today published an implementation plan for the initiative as a whole. The plan focuses primarily upon the steps necessary to complete the first databank version and benchmarking and assessment cycle. Comments upon this document are welcome.

Tuesday, April 26, 2011

Steering committee terms of reference and meeting minutes

The latest steering committee minutes are available along with a first version of their terms of reference. Forthcoming soon will be an Implementation Plan and terms of reference for sub-groups. This may all seem incredibly boring (and it generally is) but it is also absolutely necessary for the initiative to function properly if it is to be a successful multi-person, multi-institution, multi-year effort. Comments welcome as always.

Wednesday, April 6, 2011

Prototype for the databank publicly available

A very initial version of the envisaged global surface databank is available from http://www.gosic.org/GLOBAL_SURFACE_DATABANK/GBD.html. It should be stressed that this is in the very early stages of development. A full version release is not expected until early to mid 2012. This allows time to harvest additional data sources, reprocess, merge, and add provenance information so that it represents a significant delta from what has gone before in terms of both completeness and fundamental scientific value. Comments are welcome here but would be more appropriate at the new databank blog.

Tuesday, April 5, 2011

Two new initiative related blogs

There are two new more working level blogs that have been set up recently. These are more technical discussion areas than this blog. Both allow working group members to add posts and comments (moderated) are allowed from anybody else.

http://globalsurfacedatabank.blogspot.com/ covers work towards a global surface databank. If you know of data sources please head on over and provide leads in the post comments.

http://surftempbenchmarking.blogspot.com/ covers work towards a set of benchmark analogs to the databank that algorithm creators can run their algorithms on to ascertain both absolute and relative performance.

Wednesday, March 23, 2011

Data provenance and versioning task team

There is a data provenance and versioning task team for the databank effort - details are available from http://www.surfacetemperatures.org/databank/provenance-and-version-control-task-team. This task team has now met twice by telephone and made some initial inroads into the problem.

More generally the main surfacetemperatures.org site continues to be updated with progress as it happens. Other commitments have meant many of these have been left not noted here. My apologies.

Friday, January 28, 2011

Data bank task team on data rescue set up

Further details can be found here.

For anyone wondering how they can help to improve the data holdings right now, you might want to take a look at either http://www.data-rescue-at-home.org/ for upper air and land surface records, or http://www.oldweather.org/ for world war 1 UK ship records. It is hoped that many of the substantial land data holdings currently available only in image / hard copy form can eventually be digitized by citizen scientists over the internet.

Second teleconference of steering committee

Notes are posted here. Comments welcome.

Tuesday, November 30, 2010

Wednesday, November 3, 2010

Databank working group - first teleconference

The databank working group recently held their first teleconference. Notes arising from this are available here.

Comments welcome, but there will be patchy moderation, if any, until 15th, please be patient.