Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Wednesday, 27 May 2015

Where the Bradford Factor goes Bad



















This article was prompted by a story in “The Guardian”:

“You are a healthcare professional working with elderly people. You really have no business being at work because your throat is sore and inflamed; you feel febrile and have a cough. In short, you are going down with something which is, at best, a cold - but may be flu.”

“Exposure to the illness would be potentially lethal for your frail patients, but sick leave could result in you facing a formal occupational health or disciplinary hearing. You are on the horns of a dilemma already facing many health workers this winter, but one which will soon face many more.”

““This is the "Bradford factor", a system for differentiating longer, infrequent staff absences from those which are more frequent, but of shorter duration. The assumption behind it is that high frequency, short duration absences are more problematic in certain types of organisation, are symptomatic of different problems and require separate identification in order to be addressed.” (1)


How does this work?

The Bradford Index equals the number of spells of absence in the last 12 months squared, multiplied by the number off days off - (SxSxD)

So by way of example:

one absence of 14 days is 14 points (i.e. 1 x 1 x 14)
seven absences of two days each is 686 points (i.e. 7 x 7 x 14)
14 absences of one day each is 2,744 points (i.e. 14 x 14 x 14)

It is a common tool in use by Human Resources department, and apparently originated during the 1980s with a connection to the Bradford University School of Management. However, I have been unable to track down any provenance for it.

In fact, another individual who investigated found this was a Dr Geoff Helliwell who noted:

“IDS first referred to the Bradford factor or formula in Study 365 published in July 1986. That publication included a reference to the use of `Bradford' factor point scores by May & Baker (now Rhone Poulenc Rorer). Contacts at the company say that this method of calculating absence rates had been introduced as a result of managers attending a series of seminars on production management arranged as a part of a Bradford University management course in the mid-1980s. Controlling absence was one of the management areas that was covered. However, as you discovered, if you contact Bradford University's School of Management they know very little about it and I have never come across any published record of its academic origins.” (11)

Insofar as it appears anywhere in literature, Dr Helliwell found that the earliest mentions were: IDS Study 365, 'Absence', July 1986, page 5, IDS Study 498, 'Controlling Absence', January 1992, pages 3/4 (successful use of the formula at Victoria Coach Station), IDS Study 556, 'Absence & Sick Pay Policies', June 1994, page 6. He notes that a description of the factor also appears in `Measuring and monitoring absence from work' published by the Institute for Employment Studies in 1995, although its origins are not given.

This is a singular deficiency, because how it was created, what rationale lay behind the squaring of number of spells of absence, and if it was ever peer reviewed in any academic journal, is not available. In fact, as far as I can ascertain, it never has been subject to academic scrutiny in any peer reviewed journal or featured in any books on statistics.

In that respect, I find myself in agreement with a report made by Robert Perrett and Miguel Martínez Lucio in June 2006, ironically from the Bradford University School of Management, the supposed origin of the formula.

“An individual who has two spells of absence of five days in duration accumulates just 40 Bradford points, whereas a colleague who has five spells of absence of two days accumulates 250 Bradford points and is reprimanded. Use of such formulas is common. Others formulas exist in an attempt to differentiate types of absence and to provide them with a scientific rationale. Many of these formulas use spurious evidence and research, and are sometimes even difficult to identify in terms of originating documentation, but this has not deterred many organisations from using them”” (2)

The study noted that it was used as a disciplinary tool targeted at employees who played the system and were absent frequently. It is known that sick leave can be abused by employees, and that this is often short term, just the odd day.But odd days can mount up.

“Overestimating the extent of abuse can result in the implementation of harsher punitive measures that can actually result in an increase in absence rather than a decline. This was witnessed in one of the case studies undertaken. Almost half (47 per cent) of respondents believed that employees with genuine illnesses had been penalised because of the absence provisions (41 per cent did not believe this to be the case and 13 per cent were unsure) furthermore, 44 per cent of respondents claimed that sickness records formed part of an employee’s appraisal. “

“Penalising genuinely ill employees, immediately following absence or later through an appraisal system, can create a range of negative repercussions for the employee and employer alike”

And this leads to significant problems like those highlighted in the Guardian article:

“For example, employees might feel compelled to go to work to avoid being reprimanded, potentially spreading diseases. This will also have a negative effect on productivity, moral, motivation, and the desire to remain employed with the organisation.”

Philip Taylor is a Professor of Work and Employment Studies and Assistant Dean International at the Strathclyde Business School at the University of Strathclyde. He has also examined the cultural shift which has taken place:

“Taylor also points to the way in which the management of sick leave over the past two decades has become more draconian with short-term illnesses penalised through the “Bradford factor”, the use of return to work interviews and discipline and dismissal triggered by particular levels of absence. The effect of this in the workplace has been to institutionalise bullying as layers of managers are forced to “cascade” pressure to workers below them to meet targets.” (3)

Part of the problem is the focus on number. If we were to examine the Bradford Factor properly, using scientific methods, then we would have to look at the following:

What is the likelihood of any short term absence being spurious?

The Bradford factor has an inbuilt assumption that short term absences are either spurious or reveal an underlying health problem. There is no means of differentiating between the two.

For example, it cannot factor in predictable short-term absences: as for example, time off every week for treatment or counselling. The employer should accommodate this if it cannot be done outside working hours.

Nor can it deal with unpredictable short-term absences which may happen often/for a variety of reasons such as an underlying medical condition or poor immune system. This may lead to the employer suggesting an employee work flexible (or zero hour’s contract) or lighten or vary the workload for a particular period.

Now it may be argued that a high Bradford score is a good measure for identifying the basic short term absence, but the problem arises what to do when you know the reasons for the absence, because the calculation now tells you nothing more than what to expect.

In fact, a more granular computation, and one which didn't rise so spectactularly, would be to take a sum of Bradford Factors of different counts, so that a count of 1,1,2 - Bradford score on 3x3x4 = 36 would instead become 2,2 and 1,2, being 2x2x2 +1x1x2 =8+2=10. This would still isolate smaller days off but would differentiate better between them.

And it also does not help with understanding how absenteeism affects the business. Compare Tom who takes seven days off in one hit to recover from an injury and has a Bradford score of seven with Cally, who takes six absences for minor illness over six days with a Bradford score of 216.

Now if it was, for example, a business which specialised in bureau work book keeping. 25 days spread out over a year is not so critical, because the employee is back on the following day. 25 days in one go might be critical, especially where payroll functions are concerned.

But because the focus is on the high score, and not the nature of the business, this is not seen. In fact, a scoring method which rated longer periods of absence – say Tom plays rugby in his spare time, and has a few spells off - would actually be more useful.

Moreover, I have seen only one case where the Bradford factor has been analysed to look at systemic factors, type of work, different workplaces, so that a normalised benchmark can be visible. If systemic factors are causing workplace illness, then it is the workplace or working practice rather than the employee which needs to be addressed.

The return to work interview is a method of getting more information about the kind of illness, more than purely the bare bones of a Bradford figure. But the problem arises here insofar as it can be conflicted. As Phil Taylor points out:

“Return to Work Interviews (RWIs) became the most frequently utilised procedure. In practice, RWIs conflate caring and welfarist intentions (soft HRM) with calculative and disciplinary motives (hard HRM), but, principally, they reflect the latter and the impact of US thinking on occupational health, which emphasises getting people back to work rather than the employment problems that might have made people sick in the first place” (4)

The conflict here is between the use of the return to work interview as caring for the employee, and using it as a means of disciplining the employee.

The Employment Law Clinic notes this:

“Remember, this is a return-to-work Interview, not a disciplinary hearing (if a disciplinary becomes appropriate, deal with this separately, and in accordance with your disciplinary procedures)”(5)

As another site notes:

“Make sure the conversation with the absent employee is clearly focused on their well-being and their return to work. Try to focus as much on what the employee can do as things they may need help with. Returning to work is an important milestone in getting life back on track, but if the employee is made to feel a ‘problem’ in some way, they will feel disheartened.” (6)

However, in practice it may be almost impossible to construct a Chinese Wall to prevent the return to work feeding into a disciplinary, even if they are held separately. This is very clearly seen in the HROnline Website:

“From a motivational point of view, there’s evidence that companies that impose RtW interviews experience a reduction in sickness absence. There are a few reasons this could be case. First, it shows that you are taking absence seriously; that your company is taking some sort of action on absence. But second, there’s a “fear factor” involved – some employees who might in the past have taken the odd day off when they simply couldn’t be bothered to come to work might be wary of being “found out” if they slip up in the interview by contradicting their original cover story.” (7)

Clearly absenteeism is an issue that needs to be addressed, but the use of the Bradford factor as somehow more than a mathematical artefact, coupled with the “fear factor” for a return to work interview, can inculcate a hermeneutic of suspicion which is not healthy for the workplace.

As someone who has worked conducting interviews notes:

“A return to work interview is where your line manager meets with you after you have been absent (sickness, unauthorised absence, sometimes even lateness). They will sit down with you and have 'an informal chat' which involves them ticking boxes on a form and asking you how you are in a kind and caring manner (i.e. asking for answers to a set of preset questions then writing everything you say down)”

“I've had to do return to work interviews in call centres and admin jobs many times over the last 10 years, I was also a line manager in an office for a while and had to do them on other people (look, I was young and I needed the money, OK?) They are part punishment, to make you feel like you have to come up with a good explanation for your actions, (like not doing your homework at school), and to reinforce the authority of your line manager as they become the embodiment of the Divine Right of The Company (a bit like scrofula).” (8)

HR Magazine has an interesting interview with James Arquette:

“James Arquette, director at absence management company, FirstCare, is not a fan of the Bradford scale. He feels it doesn't place enough emphasis on why an absence has occurred, which is what employers and managers must look at to determine how to support members of their team. "Ultimately, it is the action of employers once they have absence data that will allow them to manage absences most effectively," says Arquette. "The Bradford scale is far too reactive, whereas a policy based on sensible 'triggers' - a specific number of stress-related absences - can allow employers not only to identify a problem, but to step in proactively and provide appropriate support."

In conclusion, the Bradford factor is a blunt tool with inbuilt assumptions about patterns of absence from the workplace. The idea that its numbers are neutral and therefore unchallengeable is quite frankly nonsense. Other computational methods can provide better answers and more granular detail.

There is no theoretical basis for that assumption presented as a hypothesis and tested, with results public in a peer review journal, not is there any other comparison with statistical patterns such as Poisson distributions.

What I would expect to see is some kind of paper like “Multiple Approaches to Absenteeism Analysis”

“This piece compares eight models that may be appropriate for analyzing absence data. Specifically, this piece discusses and uses OLS [Ordinary Least Squares] regression, OLS regression with a transformed dependent variable, the Tobit model, Poisson regression, Overdispersed Poisson regression, the Negative Binomial model, Ordinal Logistic regression, and the Ordinal Probit model. A simulation methodology is employed to determine the extent to which each model is likely to produce false positives. Simulations vary with respect to the shape of the dependent variable's distribution, sample size, and the shape of the independent variables' distributions. Actual data, based on a sample of 195 manufacturing employees, is used to illustrate how these models might be used to analyze a real data set”

Another piece is “Markov chain Monte Carlo analysis of underreported count data with an application to worker absenteeism”(1996) – “A new approach for modelling under-reported Poisson counts is developed. The parameters of the model are estimated by Markov Chain Monte Carlo simulation”

These are solid pieces of statistical research using sophisticated but standard techniques for modelling absenteeism. They can determine patterns that are endemic in the population or workplace as a whole and provide benchmarks for comparison. By contrast, the Bradford factor appears in no mathematical statistical studies. The absence of the terms "false positives" or "significance testing" from any promotions of it on HR sites should flag up a red warning that this is a statistically crude method.

As Trevor Blackmore notes succinctly, the Bradford factor has a

“very dubious pedigree. Recent enquiries to Bradford University came up blank as to who proposed this statistical monstrosity. No reliable, scientific evidence as to its effectiveness over other methods. No RCTs, no factor analysis, nothing - only unreliable case study evidence comparing a couple of years data within a few companies measuring very limited outcomes.”

That, of course, doesn’t stop people using them or believe because they crunch out numbers that they are somehow “scientific” any more than it prevents people using the Stanford-Binet IQ test, the Myers-Briggs Type Indicator, or any other pseudo-statistical measuring techniques.

The promotion of the Bradford Factor has the blurb – “Absence Management Made Simple”, which should be warning enough, although simplistic would be a better term, and probably explains why it has had such wide take up by Human Resources departments and not by professional statisticians.

Links


.(2)   Preliminary results from a survey of UNISON safety representatives: Interim Report, June 2006, Bradford University School of Management

.(3)   Taylor, Philip, 2012, “Performance Management and the New Workplace Tyranny”, Report for the Scottish Trades Union Congress, available from www.stuc.org.uk

.(4)   ‘Too scared to go sick’—reformulating the research agenda on sickness absence, Phil Taylor, Ian Cunningham, Kirsty Newsome and Dora Scholario, Industrial Relations Journal 41:4, 270–288

.(5)   http://employmentlawclinic.com/attendance-and-performance/return-to-work-interviews/

.(9)   Sturman, M. C. (1996). Multiple approaches to absenteeism analysis

.(11) https://www.jiscmail.ac.uk/cgi-bin/webadmin?A2=occenvmed;3ceccf5f.00


Tuesday, 6 January 2015

Lies, Dammed Lies, and Tourism Statistics















“The number of people visiting Jersey increased last year - with arrivals from Gatwick Airport and St Malo providing the biggest boost. A total of 1,069,265 people came through the Airport and Harbour between January and November 2014, a rise of 3.7 per cent on the previous year, according to figures released by the Economic Development Department.” (JEP)

My correspondent, Adam Gardiner, is suspicious about those figures. He comments:

“The total that came through the harbour and airport DOES NOT equal the number of people visiting Jersey. Subtract the locals travelling (and returning) also the business ‘visitors’ who simply fly in and fly out - they are NOT tourists - and business travel generally which in many cases relates to multiple journeys by the same person(s) over the course of the year.”

“I have had a bee in my bonnet for years about this so called ‘visitor numbers’. The numbers in inaccurate and misleading. I daresay numbers of tourists were up - and I may even accept 3.7% but that is not on a base number of 1.07m visitors. For a start where would that stay? We don't even have that number of bed nights available and I cannot believe day-trippers from France are increased significantly - not with the state of the Euro.”

“Until they introduce a system of counting exact number of actual visitors not just travellers in and out of the island we shall be none the wiser, and be fed these drivel statistics for political expedience and to make it look as if EDD/Tourism have been doing great job. They haven’t. I just hope this new 'tourist supremo’ will be more honest and tell us as it is - out tourism industry is all but on it’s knees. As we pull put of recession I am sure things will improve quite naturally but a renaissance it will not be without some real blue sky thinking and a huge amount of work on the strategy.”

" For what is worth (and only pure guesswork) I would estimate actual visitors ie: tourists number around 600,000 - which makes more sense with regard to the accommodation the island can provide, the number of hire cars registered and obvious lack of private investment in the general leisure economy - a sure sign in itself that tourism remains pretty stagnant."

I wondered how Guernsey does this. Do they just count arrivals? They have several documents which show how they analyse the data.

http://www.gov.gg/CHttpHandler.ashx?id=50619&p=0

This is the 2011 Travel Survey Research Report, which gives an idea of their methodology. Their numbers of departures, for instance, gives a figure of 775,500 for the tear 2011, of which it is broken down into 358,700 (46.3%) from visitors leaving the Island, and 349,800 departing residents (who would return (45.1%), along with 67,000 returning visitors leaving (8.6%).

Returning visitors are those who are counted twice in passenger numbers because they visit elsewhere during their stay in Guernsey (e.g. visitor day trips to Sark, Herm or Jersey).

Those are very useful figures, and that is the kind of data we need to know regarding Jersey. But the Jersey report (for example comparing 2014 with 2013 says “Arrivals include returning residents, visitors for all purposes including day-trippers, and returning visitors e.g. visitors to the island returning from a trip off-island.” And it gives no breakdown.

The Guernsey statisticians even have a breakdown of departing residents giving the purpose of their visit away from Guernsey.

So how do they come up with such impressively detailed figures?

“The only way to accurately measure total tourism volume is by undertaking a comprehensive exit survey in order to break down (or calibrate) passenger departure figures from the Airport and Guernsey’s Harbours. This detailed information helps the Commerce & Employment Department, Guernsey Tourism, its marketing partners and other interested parties in allocating resources, planning and refining product development and marketing strategies, and acts as a benchmark to review future progress against marketing and strategic objectives.”

And the methodology explains:

“As with previous exit surveys, face-to-face interviews were conducted with departing passengers throughout 2011, with interview shifts planned to reflect passenger throughput and to cover all routes, all days of the week and all times of the day. It is very difficult to achieve a completely randomised approach when predetermining interview shifts, but the Passenger Calibration Survey used a random sampling methodology as far as possible. Interview shifts were planned to broadly represent passenger movements throughout the year, but the selection of respondents within those shifts was random, with departing passengers being interviewed immediately after checking in at the Airport and Harbours, with the next passing person/car being selected for inclusion as soon as the previous interview had finished. This provided a randomised approach to interviewee selection, while ensuring that interviewer time was used as productively as possible.”

For more recent figures, I’ve found the 2013 Travel Survey which can be seen here.
http://www.guernseytrademedia.com/files/managed/pdf/2013_exit_survey_report_q1.pdf

The surveys are actually done by Island Ark Ltd, which is a Jersey based company! Why doesn't Jersey make use of them?

Guernsey has decided that the quality of information – and there’s a lot more in those surveys than my brief selections, including charts and comparison – is much more important than mere numbers. The survey also takes the visitors place of origin, purpose of visit, length of stay, etc.

Any serious planning for tourism needs good quality information, and simple arrival numbers, while they may have been fine during the heyday of tourism in the 1960s and 1970s, now look increasingly blunt as a means of measuring data, and giving the granular demographic information which a tourism strategy needs.

Sunday, 28 August 2011

As Usual, Average Nonsense in JEP

THE average salary in Jersey increased by just 2.5 per cent during the last year - the second-lowest rise since 1995. According to figures released yesterday, a full-time worker now earns an average of £650 per week, or £33,800 per year. Although the rise is higher than the 1.1 per cent increase seen in 2009/10, it is significantly lower than the increases seen during the previous 15 years. Economic Development Minister Alan Maclean said: 'The earnings growth rate does remain low in historic terms, but that is not unexpected given that we have a weak labour market and a challenging economic climate.'

The regular report of misinformation is in the JEP again. All across the world, the measurement of national earnings have moved from the arithmetic mean, commonly known as the "average" to the median. The arithmetic mean is obtained by adding all the items up, and dividing by the number of items. The median is a statistical measure obtained by looking at the middle item in a distribution. See "Mean Versus Median" below.

For wages, the distribution is not a balanced one - the normal distribution or "bell curve" - but a highly skewed one, with a few higher wages grossly outbalancing the majority, which is why the median is a much better measure.

The States survey even mentions this - a warning not heeded by the JEP:

The level of average earnings derived from this survey is an informative indicator, particularly when comparing sub-sectors. It should be noted when interpreting these results that as a consequence of the earnings distribution being asymmetric (i.e. skewed towards higher values) the mean statistic provides a numerically greater measure of "average" earnings than the median. There are two surveys, and the 2009/10 Jersey Income Distribution Survey gives a median, so that:

the mean average weekly earnings of full-time equivalent employees (FTE) in Jersey in June 2011 was £650 per week
the median average weekly earnings of full-time equivalent employees (FTE) in Jersey in June 2011 was £520 per week

Now £520 gives a median yearly wage of £27,040, a fair bit smaller than £33,800.

The survey also notes:

The median average cannot be determined from the data collected for the Index of Average Earnings (IAE), since calculation of the median requires earnings at an individual level rather than at a company level. The Jersey Income Distribution Survey (IDS), which was carried out over the twelve-month period from May 2009 to May 2010, collected the necessary household and individual income information required to determine median income from earnings. The results derived from the IDS data, and presented below, include an up-rate from the survey period to June 2011 using the Jersey Index of Average Earnings.

But the main survey is terribly limited in scope:

Some 430 firms in the private sector were sent a survey questionnaire; 340 completed questionnaires were received back, representing a response rate of 79%.

Considering the number of small contractors, or individuals who run their own business, this seems like a very small number of businesses to survey, and one that itself will tend to reflect larger establishments.

I understand that Guernsey, on the other hand, can supply precise median figures for wages to a far higher degree of accuracy, although I have had trouble tracking it down. When I went to a recent talk at the Hotel de France, I was amazed to discover a figure being given for median wage. "Gobsmacked" would probably describe my reaction, having been told for years by the Jersey statistics department that it was not possible because of the nature of the survey (note they need a different survey to obtain the median which is also subject to sampling errors).

I asked the presenter from Guernsey how they managed a median, and Jersey didn't, and he told me they simply culled the information from the social security records. As far as I can remember, it was around £27,000 (figures on Guernsey's "open market"  - rental without property qualifications - are lower at a median of £18,000)

This is what he told me: there is no need for a survey, all the figures are available from Social Security on all Islander's earnings

In fact, it is massively more accurate. This works because although as in Jersey, there is an earnings ceiling, the median is the mid-point, so as long as it falls below the ceiling, it can be calculated.(see Mean versus Median below for an example). Now in Jersey:

There is an 'earning ceiling' this means that there is a cap on how much social security is paid by a person.  The earning ceiling this year is £3686 a month. So if you earn above this amount you will only pay Social Security on the first £3686 you earn, so the maximum you would pay would be £221.16 (6% of £3686). The maximum your employer would pay would be £239.59 (6.5% of £3686).

This means that the median (£520 per week = £2253 per month) is well below that, so it is quite possible to do as Guernsey does, and provide a median, not based on a survey, but based on the total workforce, with income gleaned from all the Social Security records, stripped to bare numbers, and easily number crunched in today's computers. All that is needed for median, in fact, is a sort from smallest to largest, and a count to find the mid-point.
 
Perhaps one day the JEP will have some real figures to report!

Mean versus Median

As an example, consider (as wages in thousands), the following cluster:

25,25,25,25,25,25,25,25

Now this results in Average = 25, Median = 25. (For the average - the mean - we sum, and divide by 8 in this case, for the median, we take the middle of the distribution when laid out in ascending order.)

Now let's change two of the numbers to 20.

25,25,25,25,25,25,20,20

Note that we now have Average = 23.75, Median = 25. The average has slipped down slightly.

Finally, let's just replace one of those 20s with a salary of a managing directors of a medium sized finance company (and I'm sure there are plenty with higher wages) at 120. We now have:

25,25,25,25,25,25,20,120

Average = 36.25, Median = 25

Note how rapidly the average wage has shifted. We

If we are calculating from Social Security, and there is a ceiling limit, say of 44, then the figures we work from show

25,25,25,25,25,25,20,42

But as long as the median - the middle of the list - falls below that (which it does), it can still be calculated at 25.


Links

Sunday, 12 September 2010

Accidental Statistics

JERSEY'S most dangerous roads have been identified. At least 1,320 motorists and 240 pedestrians were injured in accidents on the Island's roads between January 2006 and the end of December 2009. But according to figures compiled by the Jersey Road Safety Panel, many of the smashes happened at a small number of accident blackspots. The Island's most dangerous road is St Aubin's Road, which saw 206 accidents between 2006 and 2009. Victoria Avenue is the second-most hazardous, with 173 recorded accidents during the same period. (1)

There are lies, damm lies - and of course, road statistics. What the statistics really do not say is how "dangerous" the roads really are. The roads in question are very high capacity roads, trunk roads, that convey the bulk of the traffic in and out of St Helier. Exactly the same kind of pattern emerges in a report on dangerous roads in Scotland. A road to Argyll is identified as "Scotland's most dangerous road", but it is also clear from the report that it is a key route and therefore has a high volume of traffic. Here is the report in question:

From Argyll's point of view, the area is now the known host of Scotland;s most dangerous road. The 14-mile stretch of the A819 between Inveraray and Dalmally in Argyll has been ranked at this level, with 12 crashes causing deaths or serious injuries from 2006-2008. This score is nearly three-quarters more than the total over the previous three years. Commenting today, Jamie McGrigor says: 'The Road Safety Foundation's report is yet more evidence of the need to improve the safety on the A819, something I have called for many times in the past. 'This is a key route in Argyll, both for local residents, businesses and the many tourists and visitors to the area. Everything possible must be done to reduce the accident risk on this road.

The failing in the statistics is in not correlating traffic accidents against traffic volume. If a road has a high volume of traffic, and the number of accidents are fairly evenly distributed across all the islands roads, then those with the highest volume will emerge as the most dangerous roads. But it might in fact be the case that some smaller roads see a relatively high volume of accidents over the volume of traffic passing through them -- these then would be the true dangerous roads - because it would be clear that they had a significantly higher number of accidents than perhaps roads with numerically greater accidents but a greater volume of traffic.

This is not immediately apparent, and of course the newspaper article simply reports the statistics without trying to make that much sense of them. It reports that there are more accidents when wet weather occurs after a long dry spell and more accidents may occur on the so-called more dangerous roads at particular junctions or "blackspots". What they do not do is to make the case for his judging those roads to be the most dangerous roads in the island, despite their naming the roads as such.

How the statistics play out can be seen in a well-known fallacy which arises from presenting figures without any proper interpretation. It is well known that more accidents occur in clear weather with good visibility than when there is a thick mist obscuring the driver's vision. Based purely upon numbers therefore it would seem that it would be more dangerous to be driving when it is bright and sunny than when it is thick fog. This shows very clearly the mistakes that can be made from purely relying on numbers alone. The fallacy occurs in not spotting that fog is a relatively rare condition compared to good visibility. The key here is to divide the accident statistics by the number of hours of clear weather and the number of hours of foggy weather -- this will show that accidents in fog are disproportionately higher.

When just dealing with numbers of accidents along particular roads it is easy to be seduced into thinking that the statistics are simple -- but just as with fog distorting the real import of the figures, so also traffic volume can hide which roads are really the most dangerous roads in the Island. In this way, the article by the JEP presents not just accident blackspots, but also illustrates - unintentionally - statistical blind spots.

Other factors which should be considered are (1) weather - for example, the variation in accident figures in total on the Islands roads are partly going to be effected by weather - how many times rain comes after long dry spells, how many times roads are icy, and (2) the number of cars on Islands roads. The latter, may indeed, be responsible in part from the overall increase in the figures since 2006, as more cars will certainly mean more accidents even if the chances of accidents remain the same.

What is just as dangerous as "the most dangerous roads" are the hidden factors which may make up the accident figures, and the impression of the JEP gives - that the increase is mostly due to drivers being more careless - is not as robust a suggestion as it first seems. That is not to say that care should not be taken in poor conditions, but it does mean that the ability of drivers cannot be inferred from the figures as they stand.

Links:
(1) http://www.thisisjersey.com/2010/09/11/accident-blackspots/
(2) http://forargyll.com/2010/06/a819-inveraray-dalmally-rated-scotlands-most-dangerous-roads/

Thursday, 19 November 2009

Swine Flu

We are not vaccinating our children at the moment, and this somewhat controversial blog posting (which no doubt will get some criticism) details the reasons why.

Please note that it is NOT saying that people should not have the vaccine, or have their children vaccinated. I am simply stating the special circumstances in which our children are not.

Different circumstances would have applied to my late partner, whose heart disease was of such severity that her immune system was totally compromised and any additional impact of a vaccine, however slight, (or for that matter even a common cold) could have tipped the balance, and accordingly her doctors never suggested it to her.

Last Friday, we received a letter in the post, informing us of the need for children at a local school, which I will designate School X, year ten, to have Tamiflu at Le bas Centre on Saturday, and for those children who did not receive Tamiflu, to stay away from school until Wednesday. I think it is only fair to detail a small part of our thinking on vaccinations, by way of explanation.

Our son (who is at School X) has siblings who are both on the autistic spectrum, and he himself has had a diagnosis of aspergers. We have noted in the past, that with our children, particularly two of them, there appears to be a degree in which their immune systems are compromised, and they can react badly to vaccinations.

While the vaccination programme is safe, we do not accept that it is within the bounds of probability that it is 100% safe, and the lack of proper information about small subgroups or individuals experiencing side-effects (or for that matter the protocols for VAERS and giving patients vaccine batch numbers) suggest that the complete scientific picture is not available (studies of side effects, sample sizes, statistical measures as one would find in, for example, a peer reviewed medical publication).

This is understandable because if 99% of the population benefit from the vaccination programme, political considerations make it expedient to use the vaccine; it would be foolish not to. But that does not mean that there are not a small percentage in the population who may be at risk from side effects, and accordingly, on our judgement that our son may be in a risky population for the same, he will not be participating in the programme.

In addition, in the past, on one occasion (in 1989), a batch of vaccine given to our eldest son was withdrawn on safety grounds. When we tried to get hold of details, not only was no adverse reaction noted (which he had), but also when the hospital records finally came up with a batch number and manufacturer, it turned out that it was a different manufacturer who had provided that batch number. With such inept record keeping, unless it has improved (and I'd like to know how well batch numbers are recorded against individuals in this mass programme), I have very little faith in the ability of the authorities to keep good enough records to provide a profile of the kind of group that is at risk from adverse reactions.

Tuesday, 15 September 2009

Pay Benchmarking

I've been looking at the addendum with the Hassell Blampied Associates information on pay benchmarking which is available at:.
 
http://www.statesassembly.gov.je/documents/reports/22574-18689-892009.pdf
 
From their website I notice the blurb:
 
Because of the simple, straightforward but robust methodology adopted in comparing jobs, we believe we have the best range of market information available in the Offshore Islands to ensure that businesses can take strategic decisions about their reward policies for their staff. Our survey data takes account of the differing needs in organisations from the most senior staff to the most junior, together with the clear reports that pinpoint the market information (1)
 
I can't see any explanation of this methodology, which means there is no way of checking what they mean by this at all! The survey itself says that:
 
Forty-five private sector employers participated in the survey, providing data on 3,583 jobs, and these were drawn from a variety of sectors, including Finance, Retail, Construction and the Utilities.(2)
 
Quite how there is a chief officer comparison with anyone in those areas is beyond me, especially as there are no details of what counts as "like for like". This is a complex problem which is extremely important. As the statistician Professor Joel Best notes:
 
Statistical comparisons promise a bit more-at least two numbers that might reveal a pattern: things are getting worse; or things are worse in one place than another; or this group has it worse than that one; or even this problem is more important than some others. But comparison depends on comparability. Unless each number reflects the same definitions and the same methods of measurement-unless each number is an apple, and not something else-comparisons can be deceptive. Unless the numbers are comparable, the pattern apparently revealed through comparison may say more about the nature of the numbers than it does about the nature of the social problem (3)
 
The report gives different grades of job, but nowhere does that state what the tie up is.
 
On Talkback (13/09/2009) Senator Sarah Ferguson had a solid figure for a policeman with two years training in London and Jersey, which showed a differential of £1,000 of Jersey over London in like for like occupations with the same experience and training.  That is a clear like for like comparison. In fact, police, doctors, nurses and teachers are excluded from this survey, which compares private sector to public sector pay.
 
But where do you compare, for example cleaners when there are a variety of rates and part time jobs in the private sector? Is a clerical worker at the hospital equivalent of a doctor's surgery receptionist? How do you compare catering staff (where there are staff canteens)? Is there an equivalent of a hospital porter in the private sector?
 
With a sample, normally some degree of stratification needs to be in place, over perhaps gender and age, so that the comparison reflects the private population in employment as a whole. The accompanying table gives no indication of gender, and yet other studies have shown this is important. In the USA, for example, studies noted that:
 
Men earned about the same, or less, in state and local government as comparable men earned in the private sector, but women earned more in state and local governments than in private firms.(4)
 
This means that a sample which is not stratified, and which has more women in the public sector sample than the private sector sample might well be skewed to show the employees in general in the public sector are paid more, whereas this may reflect the prevalence of better paid jobs in the public sector, where salary is awarded on the basis of grades, while in the private sector it is well known that women often receive less pay than men in comparable jobs.
 
Moreover some studies have noted that:
 
Marital status was an important influence in pay rates in the private sector, but not the public sector, and was more important for men than for women.(6)
 
Another area is part-time work, where a UK survey showed that:
 
Part-time workers received higher pay in the private sector. (6)
 
There is no breakdown of part-time employees in the table, nor is there any breakdown of the type of work, just the grade of work. And yet this too can be significant:
 
For some occupations, such as accountant, attorney, engineer, and personnel supervisor/manager, the data appear to support pay compression: the private sector pay advantage indeed rises with grade level. But for other occupations, such as computer programmer, computer systems analyst, and computer systems analyst supervisor/manager, comparisons do not show any wage compression.(6)
 
Because the table just shows the grade, and not the kind of occupation, it is impossible to see if the differentials between public and private sector are across the board in a grade, or whether part-time, female workers, different occupations may all give different results, some being more highly paid than the equivalent in the private sector, some less so.
 
Until the mid-1980s, all studies implicitly assumed that workers chose whether to work in the public or the private sector without considering pay differences between the sectors. A variety of studies since then have argued that differing pay structures attract different types of workers (e.g., if the public sector pays minorities better than the private sector does, it will attract more minority applicants) and that this will in turn affect the pay differences observed.(4)
 
This, of course, also counts for legal positions like that of the Solicitor General, Attorney-General, Deputy Bailiff and Bailiff. A lawyer may earn more in the private sector, but if they want both to be in public service, and also perhaps (because they are human, after all - and why not!), enjoy the limelight and unpaid kudos that this involves, perhaps with a knighthood, the opportunity to meet Royalty etc, the choice may not be entirely pay related.
 
Without some detail, the nice table we are given looks very solid and professional, but may simply be reflecting in part a social disparity between pay. We just do not know how the comparisons of like for like are made, and if we are in fact looking at apples and pears, both fruit, but only marginally comparable.
 
The words "simple, straightforward but robust methodology" used by the survey bring to mind the deficiencies that are often noted in surveys of this kind. The lack of transparency means we cannot tell what this is, but here is an example of how other surveys have been conducted with what one person may describe as "simple" but another as "crude", like a blunt instrument:
 
The crude measures that are used to establish comparability of individuals provide for only gross equivalence. The measures of work experience are often rough estimates that cannot separate unemployment or time out of the labor force from paid employment; they never distinguish between related and unrelated employment. Education typically is assumed to have a log-linear impact on salary, which builds in the assumptions that an additional year of education has the same percentage impact on salary whether it is at the high school, college, or postgraduate level, and that a bachelor's degree in literature increases salary as much as one in engineering does.(4)
 
The size of the firms in the sample may also be significant. Again other critical studies of comparisons have shown that:
 
Most studies do not control for the size of the employing establishment, although large companies typically pay better than small ones.(4)
 
Once more, the lack of any transparency in the published results means that it is impossible to see the size of the companies in the survey of private firms. And yet, as Joel Bests notes, this is extremely important:
 
Whenever there is disagreement about the statistical evidence, it is possible to look more closely, to discover how different measurement choices, different definitions, or other factors can explain the disparities. But, of course, this can be a lot of work; few people will make the effort to examine original sources. And, even when it is possible to clarify a specific statistical disagreement, that clarification will not resolve the larger debate about the broader social issue. Again, debates over broad social issues have their roots in competing interests and different values. While advocates for different positions tend to invoke statistics as evidence to bolster their arguments, statistics in and of themselves cannot resolve these debates. This is important because we often equate numbers with "facts." Treating a number as a fact implies that it is indisputable. It should be no surprise, then, when people interested in some social problem collect relevant statistics and present them as "facts", this is a way for them to claim authority, to argue that the facts ("It's true!") support their position
 
The table provided looks very authoritative, but how were the comparisons of like for like made? How were the private sector jobs chosen? What proportion of the whole, and how representative are they? Were there any refusals in the survey? What total percentage of jobs in the private sector does 3,583 represent? How many of the lower ones were on income support?
 
Once those questions are asked, and there is no information available that I can see for them, it seems that a lot is being taken on trust regarding the "objectivity" of the survey.
 
I have only considered the statistics, and not the causal factors behind them. But that might also be worth research. Deputy Daniel Wimberley has suggested that the higher rates of pay in the survey (given its limitations) may be due to the fact that the market goes for a race to the bottom, so that there are more private sector workers at the bottom of the scale on the minimum wage. It is not clear how many of the sample were on the minimum wage, and as he pointed out, possibly costing the State as a result of claims for income support which higher paid public sector workers, still at the bottom of the pay scale, do not. This illustrates, however, how a survey on pay alone can overlook other factors which should be considered, and begs the political question over whether low scale public sector pay should be reduced given that it may lead to more payments under income support. A single focus on pay and not income support may miss this completely.
 
I would like to end by noting Joel Best's comments on the problems that beset single studies:
 
Single studies, then, can't do the job. Absolutely every study every test, every piece of research-has limitations and flaws in its methods that make it a target for legitimate criticism. Studies should be replicated, and they should also inspire further research  that uses different methods (with, presumably, different limitations and flaws). When replication and differing methodologies confirm the same result, confidence in that finding grows. The results of' a lone study, particularly if the research raises serious methodological concerns, should not, in most scientists' view, be treated as authoritative. Only time and further research can sort out the erroneous findings from the more reliable.(5)
 
Now I am not saying that the results have been deliberately spun, and other surveys may well come to the same conclusion about pay. I am not aiming to discredit Hassell Blampied, merely to point out that one survey, with limited transparency, and no scientific peer review, can hardly be taken as authoritative, and there are weaknesses which should be addressed.
 
All I would say is that in the studies I have reviewed, the sample data and the sampling methodology has all been transparent, and published in peer reviewed journals so that other statisticians can check the results and replicate them, or test the assumptions of equivalence and parity. The note by Hassell Blampied that they have been doing these surveys for some time is not relevant to this - they could have been prone to the same methodological flaws every time through no fault of their own, and despite their complete professionalism.

Links:
1) http://www.hassellblampied.com/Salary%20&%20Benefits%20Surveys.htm
2) http://www.statesassembly.gov.je/documents/reports/22574-18689-892009.pdf
3) Dammed Lies and Statistics, Joel Best, 2001
(4) Pay, Productivity, and the Public Sector: The Case of Electrical Engineers., Langbein, Lewis, "Journal of Public Administration Research and Theory."Vol 8. Issue: 3. , 1998
(5) More Dammed Lies and Statistics, Joel Best, 2004
(6) "The Public-Private Pay Debate: What Do the Data Show?. Michael A. Miller,  Monthly Labour Review. Vol 119. Issue: 5. ,1996

Tuesday, 9 December 2008

A Third Supermarket: Possible Flaws in Methodology of the Survey

http://www.thisisjersey.com/2008/12/08/huge-vote-for-third-store/

Huge vote for third store By Dolores Cowburn

ISLANDERS have voted overwhelmingly for a third supermarket operator to come to Jersey, according to a consumer survey released today. Commissioned by Economic Development and created by the Statistics Unit, the public survey found that 84% of people wanted more choice and seven out of ten Islanders were not happy with the current range of supermarkets in Jersey. More than 1,000 people completed the survey titled 'A Third Supermarket in Jersey?', with three-quarters of those who want a third operator favouring a British chain. It is one of the biggest responses to a States survey. Only a third of people expressed concern that smaller shops may close as a result - which was one of the fears expressed by the Chamber of Commerce if another large supermarket operator was brought over.

http://www.gov.je/ChiefMinister/Statistics/News/SupermarketSurveyPage.htm

Given the size of the dataset, and using weighting to ensure all subgroups of the population are suitably represented in the analysis, we can be confident that the inferences drawn in this report robustly represent the views of Island residents.  Over two thousand households were sampled at random. These randomly selected households received a survey form through the post and were asked that the person in the household who had the next birthday (and who was aged 16 years or over), fill it in and post it back to the Statistics Unit. This method of sampling ensured that the survey randomly sampled the adult population of Jersey. The survey achieved an extremely high response rate, with 60% of sampled households filling in and returning the survey form. Such a high response rate, together with the method of sampling, ensures the sample results are both accurate and representative of the full adult Island population.

To be fair to the Statistics Unit, they then compared various Census statistics, such as age banding with the sampling done, to see how closely they matched, and found a degree - but insignificant - of younger people unrepresented in the survey. So this was a good sample.

But one with all surveys, the problems usually lies with what is not checked. One very immediate and apparent flaw - and I checked with the main document - stands out in the report. It is this - the questionnaire sent out was in English, and no mention is made of any checks to see if the minority but not insignificant Portuguese population (often the poorer members of Society) were adequately represented. Any weighting is missing here. Quite a number of this population have very poor English language skills, and answering a complex form in English, would probably be tempted to ignore it. They could have formed part of the missing 40%, the "dark figure" in the sample, and this might have produced significant differences in response to questions.

The other matter is to do with the presentation in the JEP, with its catchy "Islanders have voted..." headline. This makes the argument that if Islanders want this, they should get it, and there will be no associated problems which should be considered. Of course, it is easy to see the flaw in this - just conduct a random survey asking the question "Do you think taxes should be reduced?".

Part of this is the problem with this kind of survey itself; it only looks for immediate short-term responses, and does not see how people are actually thinking. Obviously, memory is short, because the last time a Third Supermarket was in Jersey - the original Safeway - prices did not drop significantly, as Safeway factored in not just freight costs and rental overheads, but also what the market could take. Prices for many goods in Safeway were more or less the same as in any of the other Supermarkets. Why would a Third Supermarket make a difference now? That would be a good question to ask Islanders in favour of another supermarket.

The construction of the questions also does not focus on sustainable alternatives, such as an expansion of the Farm Shop network, which has been steadily growing in the last few years, and which could be seriously damaged by another Supermarket, and which can provide produce at reasonable prices. Do you ever use a farm shop? Have you ever considered it? Do you think a third Supermarket might force farm shops out of business? Alternatives that were not really well addressed in this survey. It is well known that how questions are asked can get people to think, and deliver different outcomes, and a little section looking at this might have also been helpful.

Links:
http://www.gov.je/NR/rdonlyres/2440BD6D-B2BD-49F6-9D8F-945D14246DBF/0/SupermarketSurveyFinalreport.pdf

Monday, 29 September 2008

All I survey

http://www.thisisjersey.com/2008/09/27/what-you-think-of-the-states/

The JEP have gone to bona-fide pollsters, and the result is a survey which is "what you think of the States?"

There are two good things about this survey. It is:

a) random

Not self-selecting like the usual JEP or online poll. The problem with polls which ask people to phone in is that they are not representative of the Island as a whole, only of those people who want to reply, and may even count people twice or more. Remember the JEP phone poll on yes / no to tall buildings on the Waterfront where an automated phone system managed to clock up thousands of "yes" votes! Blog polls are the same, as they depend on checking IP addresses. Most businesses have fixed IP addresses, but the average home user is allocated an IP address (via their provider) from a pool when they connect to the internet, and can cheerfully vote each time they go on-line. I once added 30 extra votes to an on-line poll to test this.

A random poll has the advantage over this in that it takes a sample of people regardless of whether they would vote on an online poll or not. If you think "I have not been asked, and I didn't know where to find the poll", that is because the pollsters are doing their job correctly.

b) stratified by age

"Stratified" is a technical term for the JEP's own description as "posed to a sample of Islanders which was weighted according to the age profile of the electorate". That means that the pollsters get information on people's age and adjust the results in a number of possible ways to match the electorate. More accurately this poll uses a "stratified cluster" because it needs to group ages in bands.

General Comments

One way - the simplest - of doing this kind of stratified random poll is to ask the age, then discard those once you have too many for an age group. This is called "proportionate sampling" where the strata sample sizes are made proportional to the strata population sizes. For example if the first strata is made up of males, then as there are around 50% of males in the UK population, the male strata will need to represent around 50% of the total sample.
 
Another - and probably that used -is to get the results but adjust them in each group so that they match the age profile. This is the cheaper option. This is called a "disproportionate method", where the strata are not sampled according to the population sizes, but higher proportions are selected from some groups and not others. The results are then weighted to bring them to the proportions required.

As a simple example, suppose in an office complex, there are 1000 staff, 40% male, 60 % female.

We could poll 100 people in our sample, making sure there are 40 males, 60 females.

Or we could poll 100 people, and if we get (say) 45 males, 55 females, then we adjust their votes accordingly, by multiplying each part of the results by 40/45 and 60/55 to give the sample the same weighting as the original population.

Problems with the JEP survey

This works quite well, but several deficiencies are apparent.

Firstly, this is not the only way in which sampling can take place. While there may be little or no difference between male and female votes, by taking age alone, this excludes that. More importantly, it excludes different economic strata within Jersey, which is at least as likely to be representative of satisfaction with the States as age groups. I would say this is a pretty major flaw.

Secondly, this method can lead to overcompensation if there is large disproportion between the profile of the sample and the profile of the population. In our example, if they asked 10 males and 90 females, the male vote would be weighted at 40/10, the female by 90/60. This means that our 10 males are representing all the 400 males in the organisation, and even allowing for weighting it is clear this means their opinions can be disproportionate. The large the sample, the less likelihood of error - but that leads to the next problem..

Thirdly, we are not told how large the sample size is. Doubling sample size reduces sampling error by half. Along with this there is a lack of figures for sampling error, which can be expressed as a percentage range of how accurate responses will be in representing the population as a whole.

Fourthly, we are not told how the survey was conducted. Sampling is notoriously difficult to do properly because of what is termed "coverage error" which occurs because a sampling method excludes members of the population. For instance, a phone survey will not pick up on ex-directory people. The time it is taken may exclude workers, or shift-workers. Stopping people in the street only gets the people who live in that street. Mail surveys may get non-responses.

All told, the JEP survey leaves a lot out!
 

Spin and Statistics

http://www.thisisjersey.com/2008/09/27/75-increase-in-burglaries/

"75% increase in burglaries" screams the headlines.

POLICE are urging Islanders to lock doors and windows after the number of house burglaries soared by 75% this summer. A total of 150 homes have been burgled since May compared to just 86 during the same period last year.

This is a perfect example of what Darrell Huff called "How to Lie with Statistics". Of course, any burglary is bad, and I'd hate to be burgled, but let's look at the real figures.

Assume - and there are more than that - a mere 20,000 households in Jersey - flats, betsits, houses.

86 / 20,000 = 0.43 %
150 / 20,000 = 0.75%

Now recast the text as:

Burglaries sour up from 0.43% of dwellings to 0.75% of dwellings!

Not quite as impressive, but in fact more representative of the true picture. When you have small numbers, 86, 150, a jump can seem like a lot if expressed as a percentage.

Wicca, is at the moment, for example, often described as "the fastest growing religion" simply because it is expressed in percentage terms like this, taking one figure as a percentage increase on the other.  It sounds good, but all it really means is that it actually has a quite small membership, so that any increase features heavily. If expressed per head of the population, the figures rapidly reduce in significance.

It is a pity the JEP didn't ask the statisticians doing the survey (more on that later) for advice before working out the figures and coming out with a headline that is needlessly alarming.

Friday, 19 September 2008

States Strategic Plan

Some interesting statistics out of the States Strategic Plan (also known as the pre-election publicity splurge for ministers because of its interesting timing!).

http://www.gov.je/ChiefMinister/Strategic+and+Business+Planning/ProgressStatesStrategicPlan.htm

Even by their own internal assessment, Transport and Technical Services has a damming rating of 20% off track, and 50% on the amber (delays) category, mostly due to funding failures. I'd be hugely surprised in Guy de Faye retains his post even if he retains his seat, but I don't think it is entirely his fault, although his attitude doesn't help.

Some of the tick boxes seem a little out. For instance:

In 2006/7 investigate the feasibility and potential efficiency savings of providing regulatory services in partnership with Guernsey and report back to the States (CM)

That is ticked as "completed", but the outcome is rather paltry.

Various potential initiatives were raised by the CoM with Guernsey's Policy Council, none of which the Policy Council wished to pursue. This will be revisited later in 2008

If it is being "revisited", forgive me for being thick, but it doesn't really seem completed, does it?

This one is also important, and incomplete, but ticked off as complete:

In 2006 engage the relevant authorities in France, through appropriate channels, in discussions, and, in 2007 or earlier, bring forward measures to provide improved communication in relation to the nuclear activities on the Cotentin peninsula and compensation arrangements in the event of a nuclear accident (CM)

And the result:

Close links on emergencies planning and notification of incidents established between Emergency Planning Officer and authorities in Normandy; consultation with MOJ on compensation available under the Paris-Brussels conventions.

"Consultation" doesn't exactly sound as if anything is fixed regarding compensation, does it? Or is there a nice agreement lurking somewhere that I missed?

Health has several items in the "red", and with the new sunken road, the lack of progress on this one is, I feel, significantly bad:

In 2007; debate and implement an Air Quality Strategy for Jersey, including proposals for monitoring and publishing levels of local air pollution, and targets, policies and timescales for reductions in air pollution levels that reflect best practice globally (P&E) Note: This is the responsibility of the Environmental Health Team at Health & Social Services

Under "green", for ongoing and on-track, Planning have this!

Develop a viable proposal in 2006 to provide a new town park for St Helier within three to four years (P&E/TTS)

with the comment (and failure of the spell-checker in their PDF!):

Work underway to develop porposals (sic) for remediation, provision of new park and development of public car parking facilility.


and given the failure of Stuart Syvret to get anyone to look into toxic metals on the Waterfront, is it any wonder that the following is "red":

In 2007; consult on, then debate and implement, a Contaminated Land Strategy (P&E). Slippage due to competing priorities.

Social security is extremely short, and has nothing at all about supporting any kind of work schemes for mentally handicapped adults leaving the education system - who may well find that they have nothing to do except stay with their carers. It is wonderful how you can get ticks for "green" and "completed", and solve the difficult problems by just ignoring them.

Thursday, 28 August 2008

Average Nonsense



AVERAGE annual earnings in Jersey have increased by 4.3 per cent to £31,200 in the past year, according to official figures. But the increase is 0.4 per cent lower than in the previous 12 months, and the gap between pay rises and the increasing cost of living in Jersey has narrowed. The figures, released yesterday by the States Statistics Unit, shows that average weekly earnings have increased from £580 per week in June last year to £600 per week this year. But the increase in earnings was just 0.2 per cent higher than the RPI, which measures the rate of inflation, compared to 1.2 per cent higher in the previous year.

The regular report of misinformation is in the JEP again. All across the world, the measurement of national earnings have moved from the arithmetic mean, commonly known as the "average" to the median. The arithmetic mean is obtained by adding all the items up, and dividing by the number of items. The median is a statistical measure obtained by looking at the middle item in a distribution.

For wages, the distribution is not a balanced one - the normal distribution or "bell curve" - but a highly skewed one, with a few higher wages grossly outbalancing the majority, which is why the median is a much better measure.

As an example, consider (as wages in thousands), the following cluster:

25,25,25,25,25,25,25,25

Now this results in Average = 25, Median = 25. (For the average - the mean - we sum, and divide by 8 in this case, for the median, we take the middle of the distribution when laid out in ascending order.)

Now let's change two of the numbers to 20.

25,25,25,25,25,25,20,20

Note that we now have Average = 23.75, Median = 25. The average has slipped down slightly.

Finally, let's just replace one of those 20s with a salary of a managing directors of a medium sized finance company (and I'm sure there are plenty with higher wages) at 120. We now have:

25,25,25,25,25,25,20,120

Average = 36.25, Median = 25

Note how rapidly the average wage has shifted. We now have 87.5% of our sample well below the "average wage", but the median is still drawing substantially the most accurate picture.

If you want to see it graphed out with differences in the UK, go to:

(I've put a smaller version of that at the top of this blog, click on it for a bigger picture)

The UK, regarding the minimum wage, notes that:

When measuring minimum wage rates against the general level of earnings in the UK economy, we have regarded median, rather than average (mean), earnings as the more appropriate comparator. This is because of the disproportionate influence on the UK's earnings distribution of a relatively few high earners ­ which drives up the mean earnings figure.

http://www.lowpay.gov.uk/lowpay/lowpay2007/appendices4.shtml

All across Europe, Australia, America etc, the earnings statistics on wages use median rather than average. They did use average, but it was so misleading that they decided that a the median would provide a much better and more consistent measure, for the reasons outlined here.

The reason given in Jersey is that it is too difficult to get employers to supply information, yet such information is readily available from the Income Tax department, especially since the introduction of ITIS, and names and personal data could simply be stripped from the figures before number crunching. For now, take the figures given as representing the movement in wages that are appreciably higher than average, in order words, with as showing a general trend, but not one that is necessarily representative of the workforce as a whole. As a mathematician, I sigh every time these kind of statistics are trotted out, especially as they are never reported with any caveat about how misleading they are, and are no doubt swallowed whole by the general public.





Friday, 25 July 2008

Dark Figures and Number Laundering: Christian Aid's Death and Taxes

"Nine million Witches were martyred in the Burning Times."

I start with a digression, whose importance will become very clear later. It is by way of illustration as to how "dark figures" - estimates of invisible numbers - are created, and how via a process known as "number laundering", they achieve popularity and remain in circulation, even after they have been subjected to critical scrutiny and debunked.

This figure was by a German historian in the late eighteenth century who took the number of people killed in a witch hunt in his own German state and multiplied by the number of years various penal statutes existed, then reconfigured the number to correspond to the population of Europe. Professor Behringer traced the estimate of nine million victims back to wild projections made by an 18th-century anticlerical from 20 files of witch trials. The figure worked its way into 19th-century texts, was taken up by Protestant polemicists during the anti-Catholic Kulturkampf in Germany, then adopted by the early 20th-century German neopagan movement and, eventually, by anti-Christian Nazi propagandists. In the United States, the nine million figure appeared in the 1978 book "Gyn/Ecology" by the influential feminist theoretician Mary Daly, who picked it up from a 19th-century American feminist, Matilda Gage.

The modern estimate is much reduced;

For witchcraft and sorcery between 1400 and 1800, all in all, we estimate something like 50,000 legal death penalties," writes Wolfgang Behringer in "Witches and Witch-Hunts" (Polity, 2004). He estimates that perhaps twice as many received other penalties, "like banishment, fines or church penance

In fact, and almost counter to intuition, the death toll is decreasing as more and more trials are discovered.

As historian Jenny Gibbons points out:

If historians simply reported the number of executions, more deaths would obviously mean a higher death toll. But that isn't what scholars do. They can't -- they know that we're missing many court records and that several areas have never been thoroughly studied. Because of this, scholars compensate for lost records and missing data. That's why if you look at the table of estimated deaths, you'll see that the estimated death toll is about three times as high as the number of recorded executions.

Why do new trials have little impact on the death toll? Because they're replacing estimates. These newly discovered deaths aren't trials we never dreamed existed -- they come from previously unstudied areas and courts. When you "add" those new trials to the total, you also have to "subtract" the estimated deaths that scholars used to add for this same area. And since older estimates tended to be extremely high, new trial data usually ends up decreasing the death toll.

Now all that is ancient history, but it has a significant relevance to today.

Looking at Tax Evasion, the Christian Aid report looks at two methods of tax evasion -

Our figures deal with just two of the most common forms of corporate evasion. The first of these is known as 'transfer mispricing', where different parts of the companies sell goods or services to each other at manipulated prices. Again, the potential scope of this practice can be seen from the staggering fact that some 60 per cent of all world trade is now thought to take place between global corporations and their subsidiaries. The other, 'false invoicing', is where similar transactions take place between unrelated companies. We calculate, from just these two activities, that the loss of corporate taxes to the developing world is currently running at US$160bn a year (£80bn). That is more than one-and-a-half times the combined aid budgets of the whole rich world - US$103.7bn in 2007.

Our figures are derived from the work of Raymond Baker, a senior fellow at the US Center for International Policy. To arrive at his findings on transfer mispricing, he and his researchers conducted 550 interviews with heads of trading companies in 11 countries - all on condition of anonymity.

What are Christian Aid doing here? Clearly the methodology is similar to that of the counting of Witch Trials. They take the result of one study, look at how much tax is evasion is taking place from the 550 interviews in this study, and any other figures they have gleaned from public trials, and then multiplied numbers to get the larger figure. This assumes - as with the early calculations on the witch trials - that the "missing figures", what statisticians call "dark figures" - follow the same pattern on a larger scale.

Frequently we hear in reports that what is revealed - either in corporate corruption and tax evasion that comes to light - is just the tip of the iceberg, and the same kind of extrapolation as went on with the witch trials is made here. The numbers gets repeated from places ("billions in lost taxes") and by being repeated undergoes what the statistician Joel Best calls "number laundering" - it takes on a life of its own, and becomes legitimate because it is continually repeated from place to place.

A statistic's origin--perhaps simply as someone's best guess--is soon forgotten, and through repetition, the figure comes to be treated as a straightforward fact--accurate and authoritative. The trail becomes muddy, and people lose track of the estimate's original source, but they become confident that the number must be correct because it appears everywhere. It barely matters if critics challenge a number, and expose it as erroneous. Once a number is in circulation, it can live on, regardless of how thoroughly it may have been discredited. Today's improved methods of information retrieval--electronic indexes, full-text databases, and the Internet--make it easier than ever to locate statistics. Anyone who locates a number can, and quite possibly will, repeat it... Electronic storage has given us astonishing, unprecedented access to information, but many people have terrible difficulty sorting through what's available and distinguishing good information from bad. Standards for comparing and evaluating claims seem to be wanting. This is particularly true for statistics that are, after all, numbers and therefore factual, requiring no critical evaluation. Why not believe and repeat a number that everyone else uses?

In fact, the Witch trials were scattered, with clusters of intense activity, and other locations with little or no activity. The more detailed count, as Gibbons notes, reduces numbers.

Let me give you an example. Before Poland's Witch trials were counted, Bohdan Baranowski guessed that 15,000 Polish Witches died. Today, Polish historians are studying the Polish court records. They've found several hundred executions, and assume that the final death toll will be approximately 1,000. When their research is done, we will have "discovered" maybe 500 new trials. But the death toll for the Burning Times will *drop* by 14,000 because we didn't find as many trials as we expected to.

The same may well be true of the missing millions lost through corporate evasion. We simply do not know. But there is a very great danger in making use - however well intention - of figures that may well not be representative of the whole. By overstating their case in this way, Christian Aid may well ensure that in the long term, it will be seriously undermined. The corruption that comes to the surface may not be the tip of the iceberg, it may be most of the iceberg. We simply don't know. But let me finish by saying that I am sympathetic to their case, even if I think their numbers simply do not add up. As Joel best says:

When activists have generated a statistic as part of a campaign to arouse concern about some social problem, there is a tendency for them to conflate the number with the cause. Therefore, anyone who questions a statistic can be suspected of being unsympathetic to the larger claims, indifferent to the victims' suffering, and so on.


Links

Counting the Witch Trials
http://www.nytimes.com/2005/10/22/national/22beliefs.html
http://www.summerlands.com/crossroads/remembrance/_remembrance/00000082.htm

Books of the Post:
Best, Joel. Damned Lies and Statistics. Berkeley, CA: University of California Press, 2001.
Loseke, Donileen R. Thinking about Social Problems. Hawthorne, NY: Aldine de Gruyter, 1999.
Paulos, John Allen. Innumeracy. New York: Random House, 1988.

Tuesday, 8 July 2008

Lies, Damn Lies and Planning Statistics

I've been looking at the document that Planning have produced as part of their case for dismembering the Island Plan - "Review of the Island Plan to Rezone Land for Lifelong Dwellings for the over 55s and First Time Buyers - Summary of Responses" (May 2008). There are various statistics relating to questions asked such as:

Do you think land should be rezoned to help meet the needs of first time buyers (82% yes)
Do you think land should be rezoned to help to meet the needs of social rent housing for the over 55s? (64% yes)
Do you think land should be rezoned to help meet the needs of housing for the over 55s enabling home owners to downsize (69% yes)

Some very pretty pie charts are given. However...

The statistics given are from a very small number of responses - 86 written responses, and several public meetings. The comments derived from public meetings were from small numbers (100 people at any one meeting at most) and do not indicate how many of those attending agreed with the statements made.

Basically, what we have with this report is a self-selected study sample. This is very weak statistically, especially given the small numbers of responses. Self-selected samples are notorious for bias, with interest groups and activitists dominating. For example, G.R Langlois Ltd, a local builder, is one of those submitting a written response. There are not (as far as I am aware) one builder for every other 85 members of the population, so any such submission is disproportionate in its weighting. In this kind of submission, to report the views of such a narrow group of people as representative of the Island as a whole seems unbalanced in the extreme.

Statistically, a self-selective study sample can only be used to establish a hypothesis. A hypothesis established on the basis of the characteristics of a self-selected sample can be valid for the sample population, but it cannot be used to draw valid conclusions about the general population from which the self-selected or skewed sample was chosen. It needs proper testing by random sampling and larger numbers.

No valid projections can be made from the results of a statistical analysis of a non-randomly selected study sample. This is because the sample is not representative of the full population, so that projecting data beyond the sample is not justified., because the sampling error is unknown and cannot be measured.

This is what vitiates any opinion poll run by Channel Television or the Jersey Evening Post, and makes them merely "fun" . These polls do not work as a barometer of views precisely because they rely on respondents returning surveys themselves, which more often than not produces a self-selecting sample who are usually biased in one direction or another. This is why the main polling companies will seek out a representative sample ( a necessary condition for predicting the behaviour of a wider group). But Planning are seeking to base a decision on precisely the same kind of response as these opinion polls!

The advantage of probability sampling is that sampling error can be calculated. Sampling error is the degree to which a sample might differ from the population. When inferring to the population, results are reported plus or minus the sampling error. In nonprobability sampling, the degree to which the sample differs from the population remains unknown.


A "self-selecting" sample - getting people to actively respond to proposals - is simply not comparable to another sample whose members were selected at random. For this reason, the study itself, should come with a warning emphasizing that those who chose to participate may not be representative of the general population, and that the unfeasibility of obtaining a representative sample constitutes a major limitation of this study.

I would recommend that Planning contact the States Statistical Department, and ask how to set up proper stratified samples, so that they don't come out with results that, quite honestly, would be good examples in an A-Level mathematics course of how not to present statistics.

For a detailed guide to sampling bias, I would recommend for the layman

Darrell Huff's "How to Lie with Statistics", still in print, and still the easiest introduction for the non-mathematician. A snip at £5.99

http://www.amazon.co.uk/How-Lie-Statistics-Penguin-Business/dp/0140136290/ref=sr_1_1?ie=UTF8&s=books&qid=1215501549&sr=8-1


The Tiger That Isn't: Seeing Through a World of Numbers by Michael Blastland and Andrew Dilnot (
£9.09

http://www.amazon.co.uk/Tiger-That-Isnt-Through-Numbers/dp/1861978391/ref=sr_1_1?ie=UTF8&s=books&qid=1215501670&sr=1-1

Darrell Huff and Fifty Years of How to Lie with Statistics, online at:
http://www-stat.wharton.upenn.edu/~steele/Publications/PDF/SteeleSS2005.pdf

How to Lie with Statistics by Kjell Konis, online at
http://www.stats.ox.ac.uk/~konis/talks/HtLwS.pdf

And from Joel Best, Professor and Chair of Sociology and Criminal Justice at the University of Delaware, some papers online at
http://www.statlit.org/Best.htm


and also some books:

Damned Lies and Statistics: Untangling Numbers from the Media, Politicians and Activists - Joel Best
http://www.amazon.co.uk/Damned-Lies-Statistics-Untangling-Politicians/dp/0520219783/ref=pd_rhf_f_t_cs_1


More Damned Lies and Statistics: How Numbers Confuse Public Issues - Joel Best
http://www.amazon.co.uk/More-Damned-Lies-Statistics-Numbers/dp/0520238303/ref=pd_sim_b_2


And see also, on kinds of Sample
http://www.statpac.com/surveys/sampling.htm