Posts

Hoe moorddadig zijn Luxemburgers eigenlijk?

Image
Op 11 April verscheen in de krant "De Morgen" een artikel met als kop: " Meer moorden in Brussel dan in Londen en Parijs ". Het artikel behandelt het 'Global Study on Homicide 2013'-rapport, gepubliceerd door de Verenigde Naties. De journalist geeft de top 5 weer van de Europese hoofdsteden waarin het meeste aantal moorden gebeuren en zoomt dan in op de plaats van België: Helemaal bovenaan prijkt Tirana (Albanië) met 6,7 moorden per 100.000 inwoners in 2012. Tallinn (Estland), Chisinau (Moldavië), Riga (Letland) en Moskou vervolledigen de top vijf. Brussel staat op de twaalfde plaats Brussel met 2,6 moorden. In West-Europa heeft alleen Luxemburg meer moordgevallen met 3,2 per 100.000 inwoners.   Vooral de laatste zin, waarin een vergelijking wordt gemaakt met Luxemburg, is nogal ongelukkig gekozen. Dries Benoit verwees op Twitter naar een blog post van hem waarin hij, naar aanleiding van "Het Gemeente-Rapport" van Het Nieuwsblad, uitlegt waar...

Over Lampedusa, asielaanvragen, Europa en een grafiek in De Standaard

Image
Op vrijdag 25 Oktober 2013 verscheen er in "De Standaard" een artikel onder de kop "Vluchtelingen moeten het doen met beloftes". Het artikel zelf is prima, het handelt over het probleem van de vluchtelingen in Europa, dat omwille van de ramp voor Lampedusa, hoog op de Europse agenda is geraakt. De grafiek bij het artikel, echter, is niet onmiddellijk een schot in de roos te noemen. Het probleem bij deze grafiek is dat men de oppervlakte van cirkels gebruikt om verhoudingen te vergelijken, en dat is bijzonder moeilijk. Neem bijvoorbeeld het Verenigd Koninkrijk. Ongeveer de helft (14600) van de 28200 asielaanvragen wordt goedgekeurd. De oppervlakte van de rode cirkel is dan ook ongeveer de helft van de blauwe cirkel. Ik heb het eens nagerekend, en het klopt vrij aardig, maar de modale lezer zal allicht niet onmiddellijk aan die verhouding denken. Maar bon, de getallen zelf staan er netjes bij, dus ook al werkt het visueel niet goed, dan heb je toch nog de getallen ...

Managing Data Scientists

Image
With the rise of the 'Data Scientist', a lot has been said about the definition, role, qualifications and skills of the Data Scientist, and how to hire them. A somewhat neglected topic is how to manage data scientists. Indeed, data scientists, by their very nature, are hard to manage. They love to resolve problems, but those problems are not always the business problems you want them to tackle. They are ace players, but they're not always the best team players and some of them can sometimes have difficulty in dealing with (higher) management. They can have bright ideas, but they often lose interest when it comes to implementing those ideas in a profit making activity. They will find clever solutions for you, but they don't always excel in making sure that a structured process is place, let alone the administrative follow up that comes with it. Some of them were hired as 'rock-stars' and have developed an ego that goes with that... On the other hand, they are...

A small experiment with Twitter's language detection algorithm

Image
Some time a go I captured quite a lot of geo-located tweets for a spatial statistics project I'm doing. The tweets I collected were all confined to be in Belgium. One of the things I looked at was the language of tweets. As you might know, Belgium officially has three languages, Dutch, French and German. Of course, when you analyze a large set of tweets, you can't manually determine the language, on the other hand blindly relying on Twitter's language detection algorithm doesn't feel good either. That's why I set up a little experiment to assess to what extent Twitter's language detection algorithm can be trusted, in the context of  my geo-location project. I stress this because I don't have the ambition to make overall judgments on how Twitter takes care of language detection. First, let's look at the languages as determined by the Twitter language detection algorithm of the 150,000 or so tweets I collected. The barchart below shows the frequency of...

De Moivre's equation and the solar panels of Lo-Reninge

Image
A few weeks a go I saw an innocent little article  on solar panels in the Flemish quality newspaper ' De Standaard ', entitled " Niemand maakt meer zonne-energie dan inwoners Lo-Reninge ", which roughly translates to " no one produces more solar energy than the inhabitants of Lo-Reninge ". The article reports on the production of solar energy by individual households, typically produced by small installations on rooftops. The Flemish authorities support solar energy by subsidizing households who install solar panels. An important part of the subsidies is handled by issuing so called 'Green certificates' (or renewable energy certificates)  per fixed amount of  'kilowatt per hour' produced. See here for more details on solar power in Belgium. De Standaard newspaper, citing data from the Flemish Regulator of the Electricity and Gas market ( VREG ), reported on the number of these certificates issued in 2012 relative to the number of inhabitant...

An introduction to probability theory with Elvis Costello

Last week I released a paper entitled "The Generalized $S^3$-problem. A probabilistic view on Elvis Costello's Spectacular Spinning Songbook". You can find the pdf here . The paper is bit of a parody on statistical papers, so it shouldn't be taken too seriously. But at the same time it gives a very gentle introduction in some concepts of probability theory (Laplace, independence, the birthday paradox, ...). Enjoy!

Are partygoers in Belgium using more cocaine?

Image
Last week the Belgian newspaper De Morgen ran an article on drug use amongst Belgian partygoers. The headline of the article was "Partygoers use less cannabis and more cocaine" ("Minder cannabis, meer cocaïne bij feestvierders"). The graph that accompanied the article looked like this: While this is dutch, the language of drugs is universal, so I'm sure you will have no difficulty in understanding what it says. There are a couple of remarks to make on this graph: While there are small grey bars between the 3 groups, Alcohol/Cannabis, Xtc/Cocaine and LSD/GHB/Ketamine, initially I was fooled by thinking they were all using the same Y-axis. They're not, so you need to be careful to take scale into account. Secondly, at the first glance there seems to be a drop in cannabis use, but the increase in cocaine that was mentioned in the title is less clear cut (no pun intended). Thirdly, alcohol use seems to decline as well, although this is difficult to j...