Posts

Showing posts with the label Data Science

Calibration, weighting and post-stratification in audience measurement

 Read my post on substack:  Calibration, weighting and post-stratification in audience measurement

Het aandeel blanco en ongeldige stemmen bij de gemeenteraadsverkiezingen in Vlaaanderen in 2024 is gedaald, maar niet overal even sterk.

Image
Enkele weken geleden maakte de Vlaamse overheid de publicatie van de fijnmazige stemresultaten van de afgelopen lokale en provinciale verkiezingen bekend . Als datawetenschapper was ik meteen geïnteresseerd in wat deze fijnmazige resultaten juist inhielden. Wat je dan in eerste instantie vaak doet is eenvoudige data exploratie eerder dan onmiddelijk beginnen te modelleren. In eerste instantie ging mijn aandacht naar de resultaten op het niveau van telbureaus en kiesbureaus, en de mate waarin de variatie tussen telbureaus en kiesbureaus binnnen een gemeente zich verhoudt tot de variatie tussen gemeenten. Al snel viel mijn oog op het feit dat het aandeel van blanco en ongeldige stemmen overal sterk was gedaald, maar de mate waarin sterk geografisch bepaald was. Vooreerst, het feit dat het aandeel blanco en ongeldige stemmen sterk gedaald is, hoeft niet te verrassen aangezien vanaf 2024 de stemplicht in Vlaanderen werd afgeschaft. Ik merk hier meteen op dat dit niet het geval was in Bruss...

A note on observing zero successes

Image
Say that you have a sample of size $n=1000$ and you observed $S_n=100$ successes. Traditionally you would use $\hat p=\frac{S_n}{n}=\frac{100}{1000}=0.1$ as a point estimate of the population proportion $p$. From a frequentist perspective you would probably also report a confidence interval: $$p_-=\hat p - z_\alpha\sqrt{\frac{\hat p(1-\hat p)}{n}}=0.1-1.96\sqrt{\frac{0.1 \times 0.9}{1000}}=0.08140581,$$ and $$p_+=\hat p + z_\alpha\sqrt{\frac{\hat p(1-\hat p)}{n}}=0.1-1.96\sqrt{\frac{0.1 \times 0.9}{1000}}=0.1185942,$$ using $z_\alpha=1.96$ for a 95% confidence interval (Assuming that the sample fraction is small, i.e. the universe size $N$ is large relative to $n$. Also, I will not go into how such a confidence interval needs to be interpreted.). So far, so good.  Now say you have observed zero successes, i.e. $S_n=0$, and you want to apply the procedure above. To start with, you can't because it violates the non-zero sample proportion assumption.   There are some alterna...

(small) samples versus alternative (big) data sources

Image
Those of you who already have attended a meetup of the Brussels Data Science Community know that, besides excellent talks, those meetups are fun because of the traditional drinks afterwards. So after the last meetup we were on our way to a bar on the campus of the University of Brussels and I had this chat with @KrisPeeters from Dataminded. Now if you are expecting wild stories about beer and loose women (or loose men for that matter), I'm afraid I'll have to disappoint you. Instead we discussed ... sampling. Kris was questioning whether typical sample sizes market research companies work with (say in the hundreds or a few thousand at the max) still matter these days, given that we have other sources that give us much larger quantities of data. I told him everything depends on the (business) question the client has. To start with we can look at history to answer this question. In 1936 the Literary Digest poll had a sample size in the millions. But, obviously, that sample wa...

Managing Data Scientists

Image
With the rise of the 'Data Scientist', a lot has been said about the definition, role, qualifications and skills of the Data Scientist, and how to hire them. A somewhat neglected topic is how to manage data scientists. Indeed, data scientists, by their very nature, are hard to manage. They love to resolve problems, but those problems are not always the business problems you want them to tackle. They are ace players, but they're not always the best team players and some of them can sometimes have difficulty in dealing with (higher) management. They can have bright ideas, but they often lose interest when it comes to implementing those ideas in a profit making activity. They will find clever solutions for you, but they don't always excel in making sure that a structured process is place, let alone the administrative follow up that comes with it. Some of them were hired as 'rock-stars' and have developed an ego that goes with that... On the other hand, they are...

A small experiment with Twitter's language detection algorithm

Image
Some time a go I captured quite a lot of geo-located tweets for a spatial statistics project I'm doing. The tweets I collected were all confined to be in Belgium. One of the things I looked at was the language of tweets. As you might know, Belgium officially has three languages, Dutch, French and German. Of course, when you analyze a large set of tweets, you can't manually determine the language, on the other hand blindly relying on Twitter's language detection algorithm doesn't feel good either. That's why I set up a little experiment to assess to what extent Twitter's language detection algorithm can be trusted, in the context of  my geo-location project. I stress this because I don't have the ambition to make overall judgments on how Twitter takes care of language detection. First, let's look at the languages as determined by the Twitter language detection algorithm of the 150,000 or so tweets I collected. The barchart below shows the frequency of...

De Moivre's equation and the solar panels of Lo-Reninge

Image
A few weeks a go I saw an innocent little article  on solar panels in the Flemish quality newspaper ' De Standaard ', entitled " Niemand maakt meer zonne-energie dan inwoners Lo-Reninge ", which roughly translates to " no one produces more solar energy than the inhabitants of Lo-Reninge ". The article reports on the production of solar energy by individual households, typically produced by small installations on rooftops. The Flemish authorities support solar energy by subsidizing households who install solar panels. An important part of the subsidies is handled by issuing so called 'Green certificates' (or renewable energy certificates)  per fixed amount of  'kilowatt per hour' produced. See here for more details on solar power in Belgium. De Standaard newspaper, citing data from the Flemish Regulator of the Electricity and Gas market ( VREG ), reported on the number of these certificates issued in 2012 relative to the number of inhabitant...

An introduction to probability theory with Elvis Costello

Last week I released a paper entitled "The Generalized $S^3$-problem. A probabilistic view on Elvis Costello's Spectacular Spinning Songbook". You can find the pdf here . The paper is bit of a parody on statistical papers, so it shouldn't be taken too seriously. But at the same time it gives a very gentle introduction in some concepts of probability theory (Laplace, independence, the birthday paradox, ...). Enjoy!

Are partygoers in Belgium using more cocaine?

Image
Last week the Belgian newspaper De Morgen ran an article on drug use amongst Belgian partygoers. The headline of the article was "Partygoers use less cannabis and more cocaine" ("Minder cannabis, meer cocaïne bij feestvierders"). The graph that accompanied the article looked like this: While this is dutch, the language of drugs is universal, so I'm sure you will have no difficulty in understanding what it says. There are a couple of remarks to make on this graph: While there are small grey bars between the 3 groups, Alcohol/Cannabis, Xtc/Cocaine and LSD/GHB/Ketamine, initially I was fooled by thinking they were all using the same Y-axis. They're not, so you need to be careful to take scale into account. Secondly, at the first glance there seems to be a drop in cannabis use, but the increase in cocaine that was mentioned in the title is less clear cut (no pun intended). Thirdly, alcohol use seems to decline as well, although this is difficult to j...

Visualisatiefouten deren "De Morgen" niet

Image
Op woensdag 19 juni 2013 verscheen er een artikel in De Morgen met als kop " Crisis deert superrijken niet ". Eén van de twee grafieken bij het artikel verdient nadere bespreking. Ziehier de grafiek waar het over gaat: Om de tekst iets beter leesbaar te maken voor deze blog heb ik de grafiek iets aangepast: Let wel dat je rekening moet houden met de lengte verhoudingen in de eerste grafiek. Het eerste dat opvalt is dat de lengte van de twee kleinste staafdiagrammen niet in verhouding staan met de blauwe getallen (de frequenties, dus).  Voor de hoogste frequentie is er nog een excuus omdat daar een  zogenaamde schaalonderbreking wordt weergegeven (i.e. de onderbreking halverwege de staaf met de hoogste frequentie). Zoals de grafiek er nu staat had men ook een schaalonderbreking bij de 1.068.500 moeten zetten, maar aangezien de hoogte van de eerste staaf arbitrair is ten opzichte van de voorgestelde frequentie, zouden twee schaalonderbrekingen bij een grafiek met drie ...

Addendum bij "Enkele bedenkingen bij de recente "De Standaard/VRT/TNS" peiling"

Beste Tim en @_3s_, Vooreerst dank voor jullie reacties op Enkele bedenkingen bij de recente "De Standaard/VRT/TNS" peiling . Ik wil er wel meteen aan toe voegen dat het niet mijn bedoeling was om Maarten op z'n plaats te zetten, zoals Tim schrijft. Wel in tegendeel, ik vind dat Maarten intuïtief een juiste redenering had opgezet. Wat betreft m'n opmerking over de Bayesiaanse redenering van Maarten, dat was eerder als grap/compliment bedoeld. Als @_3s_ zegt dat dit niet Baysesiaans is, geloof ik hem vrij, hij is daar meer specialist in dan ik. Ik meen wel, dat in het specifieke geval van het TNS onderzoek, de journalisten gelijk hadden op te focussen op de daling die voor NVA werd geobserveerd. Uiteraard ben ook ik ervan overtuigd dat je in het algemeen ook moet kijken naar de onzekerheid die er heerst rond het vergelijkingspunt. Het maakt inderdaad uit of dat komt van de verkiezingsuitslag (geen steekproeffout, zeer kleine meetfout) of van een andere opiniepeil...

Enkele bedenkingen bij de recente "De Standaard/VRT/TNS" peiling

Ik geef geregeld commentaar op de verslaggeving over peilingen en aanverwante onderwerpen op deze blog. Bij de recente DS/VRT peiling heb ik dat niet gedaan, omdat ik al bij al vond dat de verslaggeving niet zo slecht was. Ik heb niet alle artikels gelezen, maar in het algemeen staarde m'n zich niet blind op kleine verschillen en werd de betrekkelijkheid van de resultaten vrij goed onderstreept. Tussen haakjes, Maarten Lambrechts (@maartenzam) maakte wel een aardig overzicht van de verschillende visuele weergaven van de peilingsresultaten. Ik was dus niet van plan om te regearen, maar, op populair verzoek (nu ja, enkel @janvandenbulck) toch enkele bedenkingen, met name over een twitter conversatie tussen @OmbudsDS en @maartencorten. Het uitgangspunt was de bijdrage van @OmbudsDS waarin hij schreef dat de berichtgeving over de peiling in zijn krant over het algemeen goed was. Eén van de argumenten was dat de berichtgeving zich spitste op de significante daling voor de NVA en niet...

A reaction on "On a First-name Basis with Success? Your Mom Chose Your Name Wisely."

Image
Earlier this week, the Business section of the Flemish quality newspaper 'De Standaard' reported that the shorter the first name, the higher the income (see here ). The article showed a pricture of Bill Gates, with the caption: "Was using the nickname 'Bill' the key to the success of William Henry Gates?". The newspaper was refering to research carried out by TheLadders , a "job-matching service for career-driven professionals" and reported here . Basically, they analyzed data around first names from TheLadders’ nearly 6 million members and salary level. The blog is more tongue in cheek than De Standaard article led us to believe, but the blog has found its way in social media, being liked and tweeted more than thousand times, and was caught up by the popular (and sometimes serious) press. There are, however, a few concerns with this research. Let me mention them one by one: The first concern is an obvious one: " Correlation is not caus...